From the 1 of 24 linked papers with an AI index.
24 papers
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Haodong Li, Tianfei Ren, Xiaoxiao Ma +25
The paper presents VideoCoCo, a system that generates physically consistent videos by having a coding agent produce executable Blender code that defines the scene and its dynamics,…
OpenCoF: Learning to Reason Through Video Generation
Xinyan Chen, Ziyu Guo, Renrui Zhang +2
Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reaso…
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis
Yu Qi, Haibo Zhao, Ziyu Guo +17
Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly…
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
Kaixin Zhu, Yiwen Tang, Yifan Yang +9
High-quality 3D scene reconstruction has recently advanced toward generalizable feed-forward architectures, enabling the generation of complex environments in a single forward pass…
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
Ziyu Guo, Rain Liu, Xinyan Chen +1
Visual reasoning, often interleaved with intermediate visual states, has emerged as a promising direction in the field. A straightforward approach is to directly generate images vi…
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
Dian Zheng, Manyuan Zhang, Hongyu Li +7
Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention from researchers. The main challeng…