collaborators

9 papers

cs.CV2026

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Haodong Li, Tianfei Ren, Xiaoxiao Ma +25

The paper presents VideoCoCo, a system that generates physically consistent videos by having a coding agent produce executable Blender code that defines the scene and its dynamics,…

cs.CV2026

Reconstruction Alignment Improves Unified Multimodal Models

Ji Xie, Trevor Darrell, Luke Zettlemoyer +1

Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture. However, conventional training relies on image-text pairs (or sequences) wh…

cs.CV2026

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation

Dewei Zhou, Xinyu Huang, Xun Wang +6

Generative visual models fundamentally struggle with precise spatial control. This arises from a core disconnect: models can process textual descriptions of space but cannot direct…

cs.CV2026

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Fanqing Meng, Lingxiao Du, Zijian Wu +46

Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change in…

cs.CV2026

VideoCoF: Unified Video Editing with Temporal Reasoner

Xiangpeng Yang, Ji Xie, Yiyuan Yang +4

Existing video editing methods face a critical trade-off: expert models offer precision but rely on task-specific priors like masks, hindering unification; conversely, unified temp…

cs.CV2025

3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation

Dewei Zhou, Ji Xie, Zongxin Yang +1

The increasing demand for controllable outputs in text-to-image generation has spurred advancements in multi-instance generation (MIG), allowing users to define both instance layou…