From the 1 of 12 linked papers with an AI index.
12 papers
From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting
Zizhao Chen, Ping Wei, Guang Dai +2
The paper introduces D2DF, a one‑step video object removal framework that learns to turn coarse removal drafts into high‑quality videos via privileged distillation, and adds a self…
From SRA to Self-Flow: Data Augmentation or Self-Supervision?
Dengyang Jiang, Mengmeng Wang, Harry Yang +1
Representation alignment has become an effective way to accelerate diffusion transformer training and improve generation quality. Recent self-alignment methods, such as SRA and Sel…
PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation
Bai Qicheng, Wang Ziru, Ma Teli +3
Bimanual manipulation is essential for advanced robotic systems because it offers higher efficiency and flexibility compared to single-arm configurations. However, existing approac…
SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Training
Mengmeng Wang, Dengyang Jiang, Liuzhuozheng Li +6
Denoising-based diffusion transformers, despite their strong generation performance, suffer from inefficient training convergence. Existing methods addressing this issue, such as R…
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
Zanyi Wang, Dengyang Jiang, Liuzhuozheng Li +6
Referring Video Object Segmentation (RVOS) requires segmenting specific objects in a video guided by a natural language description. The core challenge of RVOS is to anchor abstrac…
No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
Dengyang Jiang, Mengmeng Wang, Liuzhuozheng Li +6
Recent studies have demonstrated that learning a meaningful internal representation can accelerate generative training. However, existing approaches necessitate to either introduce…