works on

From the 1 of 18 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Yung-Hsu Yang, Luigi Piccinelli, Samuel Rota Bulò +7

Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This mot…

cs.CV2026

DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving

Yung-Hsu Yang, Luigi Piccinelli, Siyuan Li +8

DVPSFormer is an online architecture that jointly estimates metric depth, semantic segmentation, and instance trajectories for autonomous driving by using explicit scene discretiza…

cs.CV2025

3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection

Yung-Hsu Yang, Luigi Piccinelli, Mattia Segu +6

Monocular 3D object detection is valuable for various applications such as robotics and AR/VR. Existing methods are confined to closed-set settings, where the training and testing…

cs.CV2025

DepthSplat: Connecting Gaussian Splatting and Depth

Haofei Xu, Songyou Peng, Fangjinhua Wang +4

Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present DepthSplat to connect Gaussian splatting and depth estimation and…

cs.CV2025

ARKit LabelMaker: A New Scale for Indoor 3D Scene Understanding

Guangda Ji, Silvan Weder, Francis Engelmann +2

Neural network performance scales with both model size and data volume, as shown in both language and image processing. This requires scaling-friendly architectures and large datas…

cs.CV2024

"Where am I?" Scene Retrieval with Language

Jiaqi Chen, Daniel Barath, Iro Armeni +2

Natural language interfaces to embodied AI are becoming more ubiquitous in our daily lives. This opens up further opportunities for language-based interaction with embodied agents,…