From the 1 of 4 linked papers with an AI index.
4 papers
Video = World + Event Stream
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +24
The paper introduces Wan-Streamer v0.3, a model that treats video as a combination of a persistent world and a dynamic event stream, enabling real-time multimodal audio‑visual inte…
Wan-Streamer v0.2: Higher Resolution, Same Latency
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises…
FreeInv: Free Lunch for Improving DDIM Inversion
Yuxiang Bao, Huijie Liu, Xun Gao +2
Naive DDIM inversion process usually suffers from a trajectory deviation issue, i.e., the latent trajectory during reconstruction deviates from the one during inversion. To allevia…
Normalizing Batch Normalization for Long-Tailed Recognition
Yuxiang Bao, Guoliang Kang, Linlin Yang +3
In real-world scenarios, the number of training samples across classes usually subjects to a long-tailed distribution. The conventionally trained network may achieve unexpected inf…