From the 1 of 7 linked papers with an AI index.
7 papers
Video = World + Event Stream
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +24
The paper introduces Wan-Streamer v0.3, a model that treats video as a combination of a persistent world and a dynamic event stream, enabling real-time multimodal audio‑visual inte…
Wan-Streamer v0.2: Higher Resolution, Same Latency
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises…
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Lianghua Huang, Zhi-Fan Wu, Wei Wang +22
We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. W…
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
Ke Liu, Jiwei Wei, Wenyu Zhang +5
With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely…
AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising
Liyuan Cui, Wentao Hu, Wenyuan Zhang +3
Real-time talking avatar generation requires low latency and minute-level temporal stability. Autoregressive (AR) forcing enables streaming inference but suffers from exposure bias…
Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis
Yikang Ding, Jiwen Liu, Wenyuan Zhang +11
Recent advances in audio-driven avatar video generation have significantly enhanced audio-visual realism. However, existing methods treat instruction conditioning merely as low-lev…