activity
20232026
most citedUniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

38 citations · 69 across the 71 of their papers we have counts for

collaborators
Showing 2024Show all

37 papers · 1 filter

cs.CV2024

Occam's LGS: An Efficient Approach for Language Gaussian Splatting

Jiahuan Cheng, Jan-Nico Zaech, Luc Van Gool +1

TL;DR: Gaussian Splatting is a widely adopted approach for 3D scene representation, offering efficient, high-quality reconstruction and rendering. A key reason for its success is t…

cs.CV20241 cited

Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description

Anna-Maria Halacheva, Yang Miao, Jan-Nico Zaech +3

3D scene understanding is a long-standing challenge in computer vision and a key component in enabling mixed reality, wearable computing, and embodied AI. Providing a solution to t…

cs.CV2024

Understanding Museum Exhibits using Vision-Language Reasoning

Ada-Astrid Balauca, Sanjana Garai, Stefan Balauca +8

Museums serve as repositories of cultural heritage and historical artifacts from diverse epochs, civilizations, and regions, preserving well-documented collections that encapsulate…

cs.CV2024

Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding

Sombit Dey, Ozan Unal, Christos Sakaridis +1

3D visual grounding consists of identifying the instance in a 3D scene which is referred by an accompanying language description. While several architectures have been proposed wit…

cs.CV2024

ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives

Yuqian Fu, Runze Wang, Bin Ren +6

Bridging the gap between ego-centric and exo-centric views has been a long-standing question in computer vision. In this paper, we focus on the emerging Ego-Exo object corresponden…

cs.CV2024

InTraGen: Trajectory-controlled Video Generation for Object Interactions

Zuhao Liu, Aleksandar Yanev, Ahmad Mahmood +7

Advances in video generation have significantly improved the realism and quality of created scenes. This has fueled interest in developing intuitive tools that let users leverage v…