4 papers
SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance
Pengyiang Liu, Zhongyue Shi, Hongye Hao +7
Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video understanding evaluation across m…
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
Zitian Wang, Yue Liao, Kang Rong +3
Preference alignment has emerged as an effective strategy to enhance the performance of Multimodal Large Language Models (MLLMs) following supervised fine-tuning. While existing pr…
MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection
Zitian Wang, Zehao Huang, Yulu Gao +2
The rise of autonomous vehicles has significantly increased the demand for robust 3D object detection systems. While cameras and LiDAR sensors each offer unique advantages--cameras…
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Jie Yang, Feipeng Ma, Zitian Wang +4
Building on the success of text-based reasoning models like DeepSeek-R1, extending these capabilities to multimodal reasoning holds great promise. While recent works have attempted…