8 papers · 1 filter
DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale
Jialiang Huang, Hongxuan Tang, Jingchang Chen +128
Large-scale agentic training and evaluation with large language models (LLMs) rely on isolated, stateful execution environments in which models inspect repositories, invoke tools,…
Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation
Yang Ding, Haoran Yu, Xin Ma +8
Human-centric audio-visual generation spans several closely related tasks: animating a person from driving speech, jointly generating speech and video from a voice reference, and s…
Vorch-Omni: Multi-Task Orchestration of Sight and Sound
Vorch Team, Xiaoyu Chen, Yang Ding +25
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented ta…
Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
Lisai Zhang, Yidi Wu, Qi Liu +7
Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated…
Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
Yaole Wang, Xiaoyu Chen, Xin Ma +5
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing me…
Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics
Shujia Li, Jianshu Hu, Haiyu Zhang +5
Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current data-driven approaches primar…