Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
Xuecheng Wu, Jiaxing Liu, Danlei Huang +8
Visual-Interleaved Chain-of-Thought (VI-CoT) enables Multi-modal Large Language Models (MLLMs) to continually update their understanding and decision space based on step-wise inter…
cs.CV2025
LongCat-Video Technical Report
Meituan LongCat Team, Xunliang Cai, Qilong Huang +8
Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational vid…