works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators

8 papers

cs.CV2026

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory

Yanbo Ding, Zhizhi Guo, Quanyue Song +4

Ripple is a system for real-time joint audio‑video generation that uses a cross‑modal recurrent memory to keep long‑term context while streaming, achieving low latency and coherent…

cs.CV2026

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

Quanyue Song, Yishan He, Yanbo Ding +4

Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for…

cs.CV2026

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

Weijie Wang, Xiaoxuan He, Youping Gu +9

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via…

cs.CV2026

MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation

Yanbo Ding, Xirui Hu, Zhizhi Guo +6

Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits…

cs.CV2026

MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation

Xirui Hu, Yanbo Ding, Jiahao Wang +4

Character image animation, which synthesizes videos of reference characters driven by pose sequences, has advanced rapidly but remains largely limited to single-human settings. Exi…

cs.CV2025

V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents

Zhengrong Yue, Shaobin Zhuang, Kunchang Li +2

Despite the recent advancement in video stylization, most existing methods struggle to render any video with complex transitions, based on an open style description of user query.…