7 papers · 1 filter
Online Video Agent Harness for Long Video Understanding
Sen Yang, Boqiang Duan, Jing Yang +7
Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense f…
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
Jesen Zhang, Ningyuan Liu, Kaitong Cai +5
Multimodal LLMs often produce fluent yet unreliable reasoning, exhibiting weak step-to-step coherence and insufficient visual grounding, largely because existing alignment approach…
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
Kaitong Cai, Jusheng Zhang, Jing Yang +4
Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy…
SirenPose: Dynamic Scene Reconstruction via Geometric Supervision
Kaitong Cai, Jensen Zhang, Jing Yang +1
We introduce SirenPose, a geometry-aware loss formulation that integrates the periodic activation properties of sinusoidal representation networks with keypoint-based geometric sup…
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
Jusheng Zhang, Kaitong Cai, Xiaoyang Guo +10
The ability to perform Chain-of-Thought (CoT) reasoning marks a major milestone for multimodal models (MMs), enabling them to solve complex visual reasoning problems. Yet a critica…
3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale
Yijia Fan, Jusheng Zhang, Kaitong Cai +3
Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semantics with detailed geometric struc…