From the 2 of 60 linked papers with an AI index.
60 papers
Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs
Yung-Hsu Yang, Luigi Piccinelli, Samuel Rota Bulò +7
Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This mot…
EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking
Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang +3
Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured rep…
DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
Yung-Hsu Yang, Luigi Piccinelli, Siyuan Li +8
DVPSFormer is an online architecture that jointly estimates metric depth, semantic segmentation, and instance trajectories for autonomous driving by using explicit scene discretiza…
Head Avatars with Dynamic Explicit Hair
Vanessa Sklyarova, Haonan Chen, Berna Kabadayi +8
We present DynHair, a novel method for tracking and modeling dynamic hair for human head avatars. From video input, we reconstruct a dynamic head avatar with an explicit strand-bas…
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
Xinggang Hu, Chenyangguang Zhang, Alexandros Delitzas +4
The paper introduces a method to build detailed, hierarchical functional 3D scene graphs for indoor environments, using open‑vocabulary visual grounding and temporal graph optimiza…
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
Hongze Wang, Boyang Sun, Jiaxu Xing +5
Object-Goal Navigation (ObjectNav) is a key capability for deploying mobile robots in everyday environments such as homes, schools, and workplaces. In this task, an agent must loca…