activity
20242026
collaborators

17 papers

cs.CV2026

ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation

Xiao Luo, Mingyang Du, Xin Zhou +5

The paper introduces ROAD, a framework that transfers semantic and structural knowledge from discriminative 3D foundation models into diffusion transformers for 3D shape generation…

cs.CV2026

Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency

Zihan Su, Teng Hu, Jiangning Zhang +4

The paper introduces Cycle-World, a framework that uses reverse‑prediction cycle consistency to reduce error accumulation in long‑horizon video generation, improving temporal consi…

cs.CV2026

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

Yinan Chen, Chuming Lin, Zhennan Chen +12

While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets and benchmarks. To bridge t…

cs.CV2026

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data

Teng Hu, Mingchun Lu, Yating Wang +6

Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a sin…

cs.CV2026

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

Yuheng Chen, Teng Hu, Yuji Wang +3

Identity-preserving video generation (IPVG) aims to synthesize high-fidelity videos that follow text prompts while faithfully preserving a reference identity. Despite recent progre…

cs.CV2026

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

Shijun Shi, Jing Xu, Zhihang Li +5

Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligned reference-pose pairs with matc…