5 papers
LangLoc: "Tell Me What You See"
Shaurya Kishore Panwar, Roham Zendehdel Nobari, Shirley Feng Yi Lau +4
We tackle fine-grained indoor localization from natural language: given a free-form description of one's surroundings, estimate the observer's 2D position and heading within a know…
From Frames to Temporal Graphs: In-Context Egocentric Action Recognition with Vision-Language Models
Bessie Dominguez-Dager, Francisco Gomez-Donoso, Miguel Cazorla +3
Action reasoning in egocentric video requires capturing fine-grained transitions of hand-object interactions, a task where general-purpose Vision-Language Models (VLMs) often strug…
ZipSplat: Fewer Gaussians, Better Splats
Alexander Veicht, Sunghwan Hong, Dániel Baráth +1
Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict one Gaussian per input pixel,…
Z-FLoc: Zero-Shot Floorplan Localization via Geometric Primitives
Ayumi Umemura, Toshinori Kuwahara, Marc Pollefeys +1
Visual localization -- estimating a camera pose within a pre-existing map -- is a fundamental problem in computer vision. Floorplans are an attractive map representation: they are…
Gravity-aligned Rotation Averaging with Circular Regression
Linfei Pan, Marc Pollefeys, Dániel Baráth
Reconstructing a 3D scene from unordered images is pivotal in computer vision and robotics, with applications spanning crowd-sourced mapping and beyond. While global Structure-from…