8 papers · 1 filter
VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes
Jingru Chen, Yiming Liu, Mingtao Chen +5
Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imp…
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
Chenfeng Wang, Wei He, Xuhan Zhu +10
In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…
Visually-grounded Humanoid Agents
Hang Ye, Xiaoxuan Ma, Fan Lu +3
Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privil…
Vision-Centric Activation and Coordination for Multimodal Large Language Models
Yunnan Wang, Fan Lu, Kecheng Zheng +4
Multimodal large language models (MLLMs) integrate image features from visual encoders with LLMs, demonstrating advanced comprehension capabilities. However, mainstream MLLMs are s…
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
Fan Lu, Wei Wu, Kecheng Zheng +7
Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have dev…
Learning Visual Generative Priors without Text
Shuailei Ma, Kecheng Zheng, Ying Wei +7
Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling up expensive. We argue that gra…