works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

Haopeng Li, Yitong Li, Junsong Chen +8

Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse atten…

cs.CV2026

Motion4Motion: Motion Transfer Across Subjects at Inference

Ling-Hao Chen, Zixin Yin, Duomin Wang +2

The paper introduces Motion4Motion, a training‑free framework that transfers motion between videos by modeling motion flow instead of relying on predefined skeletons, enabling tran…

cs.CV2026

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

Peiwen Zhang, Yufan Deng, Shangkun Sun +11

Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators and robot-specific data fine-tuned models…

cs.CV2026

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

Juncheng Ma, Jianxin Bi, Yufan Deng +19

Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remai…

cs.CV2026

LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence

Zixin Yin, Xili Dai, Duomin Wang +4

The reliance on implicit point matching via attention has become a core bottleneck in drag-based editing, resulting in a fundamental compromise on weakened inversion strength and c…

cs.GR2026

Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer

Zixin Yin, Xili Dai, Ling-Hao Chen +7

Text-guided color editing in images and videos is a fundamental yet unsolved problem, requiring fine-grained manipulation of color attributes, including albedo, light source color,…