6 papers
Vorch-Omni: Multi-Task Orchestration of Sight and Sound
Vorch Team, Xiaoyu Chen, Yang Ding +25
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented ta…
Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
Lisai Zhang, Yidi Wu, Qi Liu +7
Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated…
Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
Yaole Wang, Xiaoyu Chen, Xin Ma +5
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing me…
Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics
Shujia Li, Jianshu Hu, Haiyu Zhang +5
Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current data-driven approaches primar…
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
Haomin Wang, Jinhui Yin, Qi Wei +12
General SVG modeling remains challenging due to fragmented datasets, limited transferability of methods across tasks, and the difficulty of handling structural complexity. In respo…
MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost
Sen Xing, Muyan Zhong, Zeqiang Lai +5
In this work, we explore a cost-effective framework for multilingual image generation. We find that, unlike models tuned on high-quality images with multilingual annotations, lever…