28 papers
SeeSE3: Emergence of 3D Space in Vision Features
Caroline Chen, Sayna Ebrahimi, Fedor Kitashov +4
The paper examines whether vision foundation models implicitly encode the geometry of 3D Euclidean space, introducing probes such as a mutual neighborhood metric and a Poincaré Ada…
The Computational Basis of Confidence in Large Language Models
Dharshan Kumaran, Viorica Patraucean, Maks Ovsjanikov +3
The paper investigates what the confidence signal in large language models actually represents, showing that answer logits often act as monotonic readouts of a latent decision vari…
Gen4U: Unifying Video Generation and Understanding via Diffusion
Michael King, Aravindh Mahendran, Matthew Koichi Grimes +5
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models ov…
Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
Maysam Behmanesh, Erkan Turan, Maks Ovsjanikov
Graph alignment, the problem of identifying corresponding nodes across multiple graphs, is fundamental to numerous applications. Most existing unsupervised methods embed node featu…
The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning
Theo Uscidda, Marta Tintore Gazulla, Maks Ovsjanikov +2
Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a kno…
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
Rishabh Kabra, Maks Ovsjanikov, Drew A. Hudson +5
Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features are poorly aligned across different vis…