2 papers
cs.CV2026
SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
Yunnan Wang, Kecheng Zheng, Jianyuan Wang +8
The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal info…
cs.CV2025
Vision-Centric Activation and Coordination for Multimodal Large Language Models
Yunnan Wang, Fan Lu, Kecheng Zheng +4
Multimodal large language models (MLLMs) integrate image features from visual encoders with LLMs, demonstrating advanced comprehension capabilities. However, mainstream MLLMs are s…