18 papers
Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs
Yung-Hsu Yang, Luigi Piccinelli, Samuel Rota Bulò +7
Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This mot…
LIME: Learning Intent-aware Camera Motion from Egocentric Video
Boyang Sun, Jiajie Li, Yung-Hsu Yang +6
Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent. While vis…
TORA: Topological Representation Alignment for 3D Shape Assembly
Nahyuk Lee, Zhiang Chen, Marc Pollefeys +1
Flow-matching methods for 3D shape assembly learn point-wise velocity fields that transport parts toward assembled configurations, yet they receive no explicit guidance about which…
Geometric Action Model for Robot Policy Learning
Jisang Han, Seonghu Jeon, Jaewoo Jung +7
Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-acti…
PROSE: Training-Free Egocentric Scene Registration with Vision-Language Models
Zhiang Chen, Nahyuk Lee, Boyang Sun +4
Registering two captures of the same indoor space taken at different times underpins persistent spatial memory for robots and AR systems, yet the realistic version of this task is…
ZipSplat: Fewer Gaussians, Better Splats
Alexander Veicht, Sunghwan Hong, Dániel Baráth +1
Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict one Gaussian per input pixel,…