collaborators

5 papers

cs.CV2026

LangLoc: "Tell Me What You See"

Shaurya Kishore Panwar, Roham Zendehdel Nobari, Shirley Feng Yi Lau +4

We tackle fine-grained indoor localization from natural language: given a free-form description of one's surroundings, estimate the observer's 2D position and heading within a know…

cs.CV2026

From Frames to Temporal Graphs: In-Context Egocentric Action Recognition with Vision-Language Models

Bessie Dominguez-Dager, Francisco Gomez-Donoso, Miguel Cazorla +3

Action reasoning in egocentric video requires capturing fine-grained transitions of hand-object interactions, a task where general-purpose Vision-Language Models (VLMs) often strug…

cs.CV2026

ZipSplat: Fewer Gaussians, Better Splats

Alexander Veicht, Sunghwan Hong, Dániel Baráth +1

Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict one Gaussian per input pixel,…

cs.CV2026

Z-FLoc: Zero-Shot Floorplan Localization via Geometric Primitives

Ayumi Umemura, Toshinori Kuwahara, Marc Pollefeys +1

Visual localization -- estimating a camera pose within a pre-existing map -- is a fundamental problem in computer vision. Floorplans are an attractive map representation: they are…

cs.CV2024

Gravity-aligned Rotation Averaging with Circular Regression

Linfei Pan, Marc Pollefeys, Dániel Baráth

Reconstructing a 3D scene from unordered images is pivotal in computer vision and robotics, with applications spanning crowd-sourced mapping and beyond. While global Structure-from…