From the 1 of 15 linked papers with an AI index.
15 papers
Video = World + Event Stream
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +24
The paper introduces Wan-Streamer v0.3, a model that treats video as a combination of a persistent world and a dynamic event stream, enabling real-time multimodal audio‑visual inte…
Wan-Streamer v0.2: Higher Resolution, Same Latency
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises…
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Lianghua Huang, Zhi-Fan Wu, Wei Wang +22
We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. W…
OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning
Yunyang Ge, Xianyi He, Zezhong Zhang +4
Diffusion Transformers achieve strong video generation quality, but the quadratic cost of full attention limits efficiency. We introduce OSP-Next, an efficient text-to-video genera…
FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
Yunyang Ge, Xinhua Cheng, Chengshu Zhao +5
In Image-to-Video (I2V) generation, a video is created using an input image as the first-frame condition. Existing I2V methods concatenate the full information of the conditional i…
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
Zongjian Li, Zheyuan Liu, Qihui Zhang +10
Instruction-based image editing has achieved remarkable progress; however, models solely trained via supervised fine-tuning often overfit to annotated patterns, hindering their abi…