4 papers
GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding
Mayank Nautiyal, Li Ju, Andreas Hellander +2
Standard dual-encoder vision-language models that map images and text to deterministic points on a shared unit hypersphere through normalization typically expose neither \…
Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow Matching
Li Ju, Mayank Nautiyal, Andreas Hellander +2
Vision-Language Models (VLMs) are typically deterministic in nature and lack intrinsic mechanisms to quantify epistemic uncertainty, which reflects the model's lack of knowledge or…
Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
Li Ju, Max Andersson, Stina Fredriksson +4
Vision-language models (VLMs) as foundation models have significantly enhanced performance across a wide range of visual and textual tasks, without requiring large-scale training f…
Cutup and Detect: Human Fall Detection on Cutup Untrimmed Videos Using a Large Foundational Video Understanding Model
Till Grutschus, Ola Karrar, Emir Esenov +1
This work explores the performance of a large video understanding foundation model on the downstream task of human fall detection on untrimmed video and leverages a pretrained visi…