11 papers
From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting
Zizhao Chen, Ping Wei, Guang Dai +2
The paper introduces D2DF, a one‑step video object removal framework that learns to turn coarse removal drafts into high‑quality videos via privileged distillation, and adds a self…
From SRA to Self-Flow: Data Augmentation or Self-Supervision?
Dengyang Jiang, Mengmeng Wang, Harry Yang +1
Representation alignment has become an effective way to accelerate diffusion transformer training and improve generation quality. Recent self-alignment methods, such as SRA and Sel…
PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation
Bai Qicheng, Wang Ziru, Ma Teli +3
Bimanual manipulation is essential for advanced robotic systems because it offers higher efficiency and flexibility compared to single-arm configurations. However, existing approac…
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
Rongbin Tan, Fangfang Lin, Zhenlong Yuan +10
Multimodal large language models (MLLMs) have shown remarkable capability in bridging visual perception and textual reasoning, enabling zero-shot understanding across diverse indus…
Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds
Liuzhuozheng Li, Zhiyuan Zhan, Shuhong Liu +5
Practically, training diffusion models typically requires explicit time conditioning to guide the network through the denoising sampling process. Especially in deterministic method…
Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization
Zanyi Wang, Fan Li, Dengyang Jiang +4
Spatio-temporal video grounding (STVG) aims to localize queried objects within dynamic video segments. Prevailing fully-trained approaches are notoriously data-hungry. However, gat…