64 citations · 251 across the 45 of their papers we have counts for
25 papers · 1 filter
Egocentric Vision Language Planning
Zhirui Fang, Ming Yang, Weishuai Zeng +5
We explore leveraging large multi-modal models (LMMs) and text2image models to build a more general embodied agent. LMMs excel in planning long-horizon tasks over symbolic abstract…
Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
Longxiang Tang, Zhuotao Tian, Kai Li +5
This study addresses the Domain-Class Incremental Learning problem, a realistic but challenging continual learning scenario where both the domain distribution and target classes va…
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
Jiaqi Wang, Yuhang Zang, Pan Zhang +31
Detecting objects in real-world scenes is a complex task due to various challenges, including the vast range of object categories, and potential encounters with previously unknown…
Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
Ronghui Li, YuXiang Zhang, Yachao Zhang +5
We propose Lodge, a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture, a…
Video Object Segmentation with Dynamic Query Modulation
Hantao Zhou, Runze Hu, Xiu Li
Storing intermediate frame segmentations as memory for long-range context modeling, spatial-temporal memory-based methods have recently showcased impressive results in semi-supervi…
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts
Yue Ma, Yingqing He, Hongfa Wang +8
Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and t…