collaborators

8 papers

cs.CV2026

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Vorch Team, Xiaoyu Chen, Yang Ding +25

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented ta…

cs.CV2026

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

Lisai Zhang, Yidi Wu, Qi Liu +7

Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated…

cs.CV2026

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

Yaole Wang, Xiaoyu Chen, Xin Ma +5

Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing me…

cs.CV2026

Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics

Shujia Li, Jianshu Hu, Haiyu Zhang +5

Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current data-driven approaches primar…

cs.CV2026

TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation

Qi Liu, Gang Yue, Mingyu Yin +7

Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separate, task-specific models. Bu…

cs.CL2026

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

DeepSeek-AI, Anyi Xu, Bangcai Lin +315

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…