8 papers
Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
Pascal Benschop, Justin Dauwels, Jan van Gemert
Spatial reasoning in vision language models (VLMs) remains fragile when semantics hinge on subtle temporal or geometric cues. We introduce a synthetic benchmark that probes two com…
Evaluation of Vision-LLMs in Surveillance Video
Pascal Benschop, Cristian Meo, Justin Dauwels +1
The widespread use of cameras in our society has created an overwhelming amount of video data, far exceeding the capacity for human monitoring. This presents a critical challenge f…
BlockGPT: Spatio-Temporal Modelling of Rainfall via Frame-Level Autoregression
Cristian Meo, Varun Sarathchandran, Avijit Majhi +6
Predicting precipitation maps is a highly complex spatiotemporal modeling task, critical for mitigating the impacts of extreme weather events. Short-term precipitation forecasting,…
Assessing the Geographic Generalization and Physical Consistency of Generative Models for Climate Downscaling
Carlo Saccardi, Maximilian Pierzyna, Haitz Sáez de Ocáriz Borde +6
Kilometer-scale weather data is crucial for real-world applications but remains computationally intensive to produce using traditional weather simulations. An emerging solution is…
Compositional Scene Understanding through Inverse Generative Modeling
Yanbo Wang, Justin Dauwels, Yilun Du
Generative models have demonstrated remarkable abilities in generating high-fidelity visual content. In this work, we explore how generative models can further be used not only to…
-TCVAE: On the relationship between Disentanglement and Diversity
Cristian Meo, Louis Mahon, Anirudh Goyal +1
While disentangled representations have shown promise in generative modeling and representation learning, their downstream usefulness remains debated. Recent studies re-defined dis…