collaborators

7 papers

cs.CV2026

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

Lisai Zhang, Yidi Wu, Qi Liu +7

Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated…

cs.CV2026

TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation

Qi Liu, Gang Yue, Mingyu Yin +7

Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separate, task-specific models. Bu…

cs.CV2026

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

Jingyuan Chen, Sheng Jin, Haopeng Sun +2

Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD research typically studies tasks in…

cs.CL2025

UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions

Wenkang Han, Zhixiong Zeng, Jing Huang +7

Autonomous agents for Graphical User Interfaces (GUIs) are revolutionizing human-computer interaction, yet their reliance on text-based instructions imposes limitations on accessib…

cs.AI2025

ScaleTrack: Scaling and back-tracking Automated GUI Agents

Jing Huang, Zhixiong Zeng, Wenkang Han +5

Automated GUI agents aims to facilitate user interaction by automatically performing complex tasks in digital environments, such as web, mobile, desktop devices. It receives textua…

cs.LG2025

sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging

Jingyuan Chen, Yuan Yao, Mie Anderson +5

Automatic sleep staging based on electroencephalography (EEG) and electromyography (EMG) signals is an important aspect of sleep-related research. Current sleep staging methods suf…