From the 1 of 8 linked papers with an AI index.
8 papers
Video = World + Event Stream
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +24
The paper introduces Wan-Streamer v0.3, a model that treats video as a combination of a persistent world and a dynamic event stream, enabling real-time multimodal audio‑visual inte…
Wan-Streamer v0.2: Higher Resolution, Same Latency
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises…
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Lianghua Huang, Zhi-Fan Wu, Wei Wang +22
We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. W…
RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control
Youcan Xu, Jiaxin Shi, Zhen Wang +5
Camera-controlled video-to-video (V2V) generation enables dynamic viewpoint synthesis from monocular footage, holding immense potential for interactive filmmaking and live broadcas…
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
Zichong Li, Chen Liang, Liliang Ren +3
Large language models (LLMs) increasingly operate in settings that require reliable long-context understanding, such as retrieval-augmented generation and multi-document reasoning.…
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
Hengye Lyu, Zisu Li, Yue Hong +4
Recent advances in video generation models has significantly accelerated video generation and related downstream tasks. Among these, video stylization holds important research valu…