works on

From the 1 of 20 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Yung-Hsu Yang, Luigi Piccinelli, Samuel Rota Bulò +7

Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This mot…

cs.CV2026

DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving

Yung-Hsu Yang, Luigi Piccinelli, Siyuan Li +8

DVPSFormer is an online architecture that jointly estimates metric depth, semantic segmentation, and instance trajectories for autonomous driving by using explicit scene discretiza…

cs.CV2026

OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos

Oussema Dhaouadi, Zuria Bauer, Johannes Michael Meier +3

Continuous 6-DoF pose estimation is essential for autonomous UAV operations. Yet, existing visual odometry and SLAM methods accumulate drift and yield only relative, up-to-scale tr…

cs.CV2026

LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection

Johannes Meier, Jonathan Michel, Oussema Dhaouadi +7

Real-time monocular 3D object detection remains challenging due to severe depth ambiguity, viewpoint shifts, and the high computational cost of 3D reasoning. Existing approaches ei…

cs.CV2026

PROSE: Training-Free Egocentric Scene Registration with Vision-Language Models

Zhiang Chen, Nahyuk Lee, Boyang Sun +4

Registering two captures of the same indoor space taken at different times underpins persistent spatial memory for robots and AR systems, yet the realistic version of this task is…

cs.CV2026

From Frames to Temporal Graphs: In-Context Egocentric Action Recognition with Vision-Language Models

Bessie Dominguez-Dager, Francisco Gomez-Donoso, Miguel Cazorla +3

Action reasoning in egocentric video requires capturing fine-grained transitions of hand-object interactions, a task where general-purpose Vision-Language Models (VLMs) often strug…