From the 1 of 4 linked papers with an AI index.
4 papers
Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
Yanbo Ding, Zhizhi Guo, Quanyue Song +4
Ripple is a system for real-time joint audio‑video generation that uses a cross‑modal recurrent memory to keep long‑term context while streaming, achieving low latency and coherent…
OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
Quanyue Song, Yishan He, Yanbo Ding +4
Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for…
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
Quanyue Song, Yishan He, Yanfei Zhang +6
Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consis…
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
Lei Yang, Yi He, Fei Wu +1
Chinese mandarin visual speech recognition (VSR) is a task that has advanced in recent years, yet still lags behind the performance on non-tonal languages such as English. One prim…