5 papers
Rethinking Expressivity and Efficiency in Test-Time Training
Zeyun Zhong, Joya Chen, Manuel Martin +3
Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token updat…
FlowNar: Scalable Streaming Narration for Long-Form Videos
Zeyun Zhong, Manuel Martin, Chengzhi Wu +4
Recent Large Multimodal Models (LMMs), primarily designed for offline settings, are ill-suited for the dynamic requirements of streaming video. While recent online adaptations impr…
CamC2V: Context-aware Controllable Video Generation
Luis Denninger, Sina Mokhtarzadeh Azar, Juergen Gall
Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditions to guide generation. However…
Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
Sina Mokhtarzadeh Azar, Emad Bahrami, Enrico Pallotta +3
In this work, we investigate diffusion-based video prediction models, which forecast future video frames, for continuous video streams. In this context, the models observe continuo…
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos +3
Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work…