activity
20242026
collaborators

18 papers

cs.CV2026

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Yung-Hsu Yang, Luigi Piccinelli, Samuel Rota Bulò +7

Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This mot…

cs.RO2026

LIME: Learning Intent-aware Camera Motion from Egocentric Video

Boyang Sun, Jiajie Li, Yung-Hsu Yang +6

Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent. While vis…

cs.CV2026

TORA: Topological Representation Alignment for 3D Shape Assembly

Nahyuk Lee, Zhiang Chen, Marc Pollefeys +1

Flow-matching methods for 3D shape assembly learn point-wise velocity fields that transport parts toward assembled configurations, yet they receive no explicit guidance about which…

cs.RO2026

Geometric Action Model for Robot Policy Learning

Jisang Han, Seonghu Jeon, Jaewoo Jung +7

Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-acti…

cs.CV2026

PROSE: Training-Free Egocentric Scene Registration with Vision-Language Models

Zhiang Chen, Nahyuk Lee, Boyang Sun +4

Registering two captures of the same indoor space taken at different times underpins persistent spatial memory for robots and AR systems, yet the realistic version of this task is…

cs.CV2026

ZipSplat: Fewer Gaussians, Better Splats

Alexander Veicht, Sunghwan Hong, Dániel Baráth +1

Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict one Gaussian per input pixel,…