activity
20242026
collaborators

6 papers

cs.CV2026

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Vorch Team, Xiaoyu Chen, Yang Ding +25

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented ta…

cs.CV2026

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

Lisai Zhang, Yidi Wu, Qi Liu +7

Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated…

cs.CV2026

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

Yaole Wang, Xiaoyu Chen, Xin Ma +5

Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing me…

cs.CV2026

Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics

Shujia Li, Jianshu Hu, Haiyu Zhang +5

Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current data-driven approaches primar…

cs.CV2025

InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models

Haomin Wang, Jinhui Yin, Qi Wei +12

General SVG modeling remains challenging due to fragmented datasets, limited transferability of methods across tasks, and the difficulty of handling structural complexity. In respo…

cs.CL2024

MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost

Sen Xing, Muyan Zhong, Zeqiang Lai +5

In this work, we explore a cost-effective framework for multilingual image generation. We find that, unlike models tuned on high-quality images with multilingual annotations, lever…