10 citations · 16 across the 14 of their papers we have counts for
13 papers · 1 filter
TimeThink: Reasoning with Time for Video LLMs
Handong Li, Longteng Guo, Zikang Liu +8
Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promisi…
Semantic-Enriched Latent Visual Reasoning
Tianrun Xu, Yue Sun, Qixun Wang +8
Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
Longteng Guo, Yifan Wang, Pengkang Huo +4
Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning…
M-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
Jiatong Ma, Longteng Guo, Yuchen Liu +4
We present M-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multi…
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
Handong Li, Zikang Liu, Longteng Guo +10
Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception throug…
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
Tongtian Yue, Longteng Guo, Yepeng Tang +4
Despite the impressive advancements of Large Vision-Language Models (LVLMs), existing approaches suffer from a fundamental bottleneck: inefficient visual-language integration. Curr…