2 papers
cs.CV2026
SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point Tracking
Weiguang Zhao, Haoran Xu, Xingyu Miao +11
Point tracking aims to follow visual points through complex motion, occlusion, and viewpoint changes, and has advanced rapidly with modern foundation models. Yet progress toward ge…
cs.CL2025
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
Zhoutong Ye, Mingze Sun, Huan-ang Gao +9
Large multimodal models (LMMs) have demonstrated significant potential as generalists in vision-language (VL) tasks. However, adoption of LMMs in real-world tasks is hindered by th…