7 papers
Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
Lisai Zhang, Yidi Wu, Qi Liu +7
Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated…
TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation
Qi Liu, Gang Yue, Mingyu Yin +7
Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separate, task-specific models. Bu…
UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD
Jingyuan Chen, Sheng Jin, Haopeng Sun +2
Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD research typically studies tasks in…
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
Wenkang Han, Zhixiong Zeng, Jing Huang +7
Autonomous agents for Graphical User Interfaces (GUIs) are revolutionizing human-computer interaction, yet their reliance on text-based instructions imposes limitations on accessib…
ScaleTrack: Scaling and back-tracking Automated GUI Agents
Jing Huang, Zhixiong Zeng, Wenkang Han +5
Automated GUI agents aims to facilitate user interaction by automatically performing complex tasks in digital environments, such as web, mobile, desktop devices. It receives textua…
sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging
Jingyuan Chen, Yuan Yao, Mie Anderson +5
Automatic sleep staging based on electroencephalography (EEG) and electromyography (EMG) signals is an important aspect of sleep-related research. Current sleep staging methods suf…