works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CV2026

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory

Yanbo Ding, Zhizhi Guo, Quanyue Song +4

Ripple is a system for real-time joint audio‑video generation that uses a cross‑modal recurrent memory to keep long‑term context while streaming, achieving low latency and coherent…

cs.CV2026

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

Quanyue Song, Yishan He, Yanbo Ding +4

Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for…

cs.CV2026

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

Quanyue Song, Yishan He, Yanfei Zhang +6

Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consis…

cs.CV2026

Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation

Baoteng Li, Xianghao Zang, Xinran Wang +8

Text-to-Image (T2I) generation has achieved remarkable progress in recent years. Meanwhile, reinforcement learning methods, particularly those based on Group Relative Policy Optimi…

cs.AI2026

DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents

Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia +7

Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instea…

cs.CV2026

MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation

Yanbo Ding, Xirui Hu, Zhizhi Guo +6

Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits…