Code Monkey home page Code Monkey logo

ivg's Introduction

Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions

Wenxuan Wang, Yisi Zhang, Xingjian He, Yichen Yan, Zijia Zhao, Xinlong Wang, Jing Liu

Paper PDF Project Page

๐Ÿšฉ Updates

Welcome to this repository for the latest updates.

โœ… [2024.2.17] : Released our paper on arXiv.

โœ… [2024.5.16] : Our paper is officially accepted by ACL 2024.

[ ] : Released our data and baselines.

๐ŸŒ• Abstract

In this work, we take a step further to the intention-driven visual-language (V-L) understanding. To promote classic VG towards human intention interpretation, we propose a new intention-driven visual grounding (IVG) task and build a largest-scale IVG dataset named IntentionVG with free-form intention expressions. Considering that practical agents need to move and find specific targets among various scenarios to realize the grounding task, our IVG task and IntentionVG dataset have taken the crucial properties of both multi-scenario perception and egocentric view into consideration. Besides, various types of models are set up as the baselines to realize our IVG task. Extensive experiments on our IntentionVG dataset and baselines demonstrate the necessity and efficacy of our method for the V-L field. To foster future research in this direction, our newly built dataset and baselines will be publicly available.


๐ŸŒ– Intention-Driven Visual Grounding (IVG) Task

๐ŸŒ— Data Collection Engine & IntentionVG Dataset

๐ŸŒ˜ Baseline Constructions

๐ŸŒ‘ Results

๐Ÿš€ Citation

If you use our data or baseline model in your work or find it is helpful, please cite the corresponding paper:

@article{wang2024beyond,
  title={Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions},
  author={Wang, Wenxuan and Zhang, Yisi and He, Xingjian and Yan, Yichen and Zhao, Zijia and Wang, Xinlong and Liu, Jing},
  journal={arXiv preprint arXiv:2402.11265},
  year={2024}
}

๐Ÿญ Acknowledgement

This work is built on many excellent research works and open-source projects, thanks a lot to all the authors for sharing!

1.EVA-CLIP

2.Qwen-VL

3.MiniGPT-4

ivg's People

Contributors

rubics-xuan avatar

Stargazers

 avatar yptang avatar Devansh Khandekar avatar  avatar  avatar  avatar SUN, Pengzhan avatar  avatar Jihwan Park avatar  avatar  avatar Joez avatar Daqiu Shi avatar yahooo avatar Pๆก‘ avatar Tongtian Yue avatar  avatar

Watchers

 avatar  avatar

ivg's Issues

IVG dataset release

Thank you for the excellent work. Could you please let me know when the dataset might be available?

Recommend Projects

  • React photo React

    A declarative, efficient, and flexible JavaScript library for building user interfaces.

  • Vue.js photo Vue.js

    ๐Ÿ–– Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.

  • Typescript photo Typescript

    TypeScript is a superset of JavaScript that compiles to clean JavaScript output.

  • TensorFlow photo TensorFlow

    An Open Source Machine Learning Framework for Everyone

  • Django photo Django

    The Web framework for perfectionists with deadlines.

  • D3 photo D3

    Bring data to life with SVG, Canvas and HTML. ๐Ÿ“Š๐Ÿ“ˆ๐ŸŽ‰

Recommend Topics

  • javascript

    JavaScript (JS) is a lightweight interpreted programming language with first-class functions.

  • web

    Some thing interesting about web. New door for the world.

  • server

    A server is a program made to process requests and deliver data to clients.

  • Machine learning

    Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.

  • Game

    Some thing interesting about game, make everyone happy.

Recommend Org

  • Facebook photo Facebook

    We are working to build community through open source technology. NB: members must have two-factor auth.

  • Microsoft photo Microsoft

    Open source projects and samples from Microsoft.

  • Google photo Google

    Google โค๏ธ Open Source for everyone.

  • D3 photo D3

    Data-Driven Documents codes.