5 papers
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos
Weisong Liu, Haochen Wang, Kuan Gao +8
We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to co…
Towards One-to-Many Temporal Grounding
Qi Xu, Yue Tan, Shihao Chen +5
Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retrieval. Real-world scenarios, ho…
GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
Xudong Lu, Zhi Zheng, Yi Wan +11
Cross-View Geo-Localization (CVGL) focuses on identifying correspondences between images captured from distinct perspectives of the same geographical location. However, existing CV…
MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
Siyuan Li, Kai Yu, Anna Wang +7
Modeling genomic sequences faces two unsolved challenges: the information density varies widely across different regions, while there is no clearly defined minimum vocabulary unit.…
Evaluating Compositional Scene Understanding in Multimodal Generative Models
Shuhao Fu, Andrew Jun Lee, Anna Wang +4
The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential for computer vision systems to…