1 citations · 1 across the 8 of their papers we have counts for
8 papers · 1 filter
Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
Lisai Zhang, Yidi Wu, Qi Liu +7
Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated…
UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD
Jingyuan Chen, Sheng Jin, Haopeng Sun +2
Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD research typically studies tasks in…
TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation
Qi Liu, Gang Yue, Mingyu Yin +7
Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separate, task-specific models. Bu…
Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion
Jingyuan Chen, Fuchen Long, Jie An +4
The first-in-first-out (FIFO) video diffusion, built on a pre-trained text-to-video model, has recently emerged as an effective approach for tuning-free long video generation. This…
The Blessings of Unlabeled Background in Untrimmed Videos
Yuan Liu, Jingyuan Chen, Zhenfang Chen +3
Weakly-supervised Temporal Action Localization (WTAL) aims to detect the action segments with only video-level action labels in training. The key challenge is how to distinguish th…
Learning to Segment the Tail
Xinting Hu, Yi Jiang, Kaihua Tang +3
Real-world visual recognition requires handling the extreme sample imbalance in large-scale long-tailed data. We propose a "divide&conquer" strategy for the challenging LVIS task:…