10 papers
GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction
Chenrui Fan, Paolo Favaro
Camera-prompted text-to-video (T2V) models are increasingly used to synthesize virtual camera captures, such as orbiting objects or moving through static scenes. For these outputs,…
World Model Self-Distillation: Training World Models to Solve General Tasks
Sebastian Stapf, Pablo Acuaviva Huertos, Aram Davtyan +1
Pretrained video generators are promising visual world models that exhibit emergent task-solving abilities; however, their reliance on detailed textual descriptions limits their di…
Faster Inference of Flow-Based Generative Models via Improved Data-Noise Coupling
Aram Davtyan, Leello Tadesse Dadi, Volkan Cevher +1
Conditional Flow Matching (CFM), a simulation-free method for training continuous normalizing flows, provides an efficient alternative to diffusion models for key tasks like image…
Communication-Inspired Tokenization for Structured Image Representations
Aram Davtyan, Yusuf Sahin, Yasaman Haghighi +4
Discrete image tokenizers have emerged as a key component of modern vision and multimodal systems, providing a sequential interface for transformer-based architectures. However, mo…
Rethinking Visual Intelligence: Insights from Video Pretraining
Pablo Acuaviva, Aram Davtyan, Mariam Hassan +4
Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This su…
KOALA++: Efficient Kalman-Based Optimization with Gradient-Covariance Products
Zixuan Xia, Aram Davtyan, Paolo Favaro
We propose KOALA++, a scalable Kalman-based optimization algorithm that explicitly models structured gradient uncertainty in neural network training. Unlike second-order methods, w…