audio-video synthesis 1cross-modal memory 1long-form generation 1real-time generation 1streaming models 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
Yanbo Ding, Zhizhi Guo, Quanyue Song +4
Ripple is a system for real-time joint audio‑video generation that uses a cross‑modal recurrent memory to keep long‑term context while streaming, achieving low latency and coherent…
cs.CV2026
MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation
Yanbo Ding, Xirui Hu, Zhizhi Guo +6
Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits…
cs.CV2026
MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation
Xirui Hu, Yanbo Ding, Jiahao Wang +4
Character image animation, which synthesizes videos of reference characters driven by pose sequences, has advanced rapidly but remains largely limited to single-human settings. Exi…