4 papers
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
Qinglin Zeng, Kaitong Cai, Ruiqi Chen +2
Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independentl…
PTTA: A Pure Text-to-Animation Framework for High-Quality Creation
Ruiqi Chen, Kaitong Cai, Yijia Fan +1
Traditional animation production involves complex pipelines and significant manual labor cost. While recent video generation models such as Sora, Kling, and CogVideoX achieve impre…
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
Jusheng Zhang, Kaitong Cai, Xiaoyang Guo +10
The ability to perform Chain-of-Thought (CoT) reasoning marks a major milestone for multimodal models (MMs), enabling them to solve complex visual reasoning problems. Yet a critica…
GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning
Jusheng Zhang, Yijia Fan, Wenjun Lin +5
We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolithic models, GAM-Agent formulates the reaso…