7 papers
Vorch-Omni: Multi-Task Orchestration of Sight and Sound
Vorch Team, Xiaoyu Chen, Yang Ding +25
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented ta…
Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
Lisai Zhang, Yidi Wu, Qi Liu +7
Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated…
Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
Menglin Han, Yang Ding, Yulei Lu +6
Real-time long-form avatar audio-video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained…
Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
Yaole Wang, Xiaoyu Chen, Xin Ma +5
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing me…
Monocular Normal Estimation via Shading Sequence Estimation
Zongrui Li, Xinhua Ma, Minghui Hu +6
Monocular normal estimation aims to estimate the normal map from a single RGB image of an object under arbitrary lights. Existing methods rely on deep models to directly predict no…
Training-free Stylized Text-to-Image Generation with Fast Inference
Xin Ma, Yaohui Wang, Xinyuan Chen +2
Although diffusion models exhibit impressive generative capabilities, existing methods for stylized image generation based on these models often require textual inversion or fine-t…