5 papers · 1 filter
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
Qian Yang, Ankur Sikarwar, Huy Le +4
Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they reason in language and discard the fine-grained geometry the task requires. Thinking with i…
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
Huy Le, Nhat Chung, Tung Kieu +2
Video Scene Graph Generation (VidSGG) aims to represent dynamic visual content by detecting objects and modeling their temporal interactions as structured graphs. Prior studies typ…
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
Huy Le, Nhat Chung, Tung Kieu +2
Text-video retrieval (TVR) systems often suffer from visual-linguistic biases present in datasets, which cause pre-trained vision-language models to overlook key details. To addres…
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
Fabian Gröger, Shuo Wen, Huyen Le +1
Multimodal models have demonstrated powerful capabilities in complex tasks requiring multimodal alignment, including zero-shot classification and cross-modal retrieval. However, ex…
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
Huy Le, Tung Kieu, Anh Nguyen +1
Text-video retrieval, a prominent sub-field within the domain of multimodal information retrieval, has witnessed remarkable growth in recent years. However, existing methods assume…