3 papers
cs.CV2026
SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes
Léopold Maillard, Francis Engelmann, Tom Durand +5
Embodied AI depends on interactive 3D environments that support meaningful activities for diverse users, yet assessing their functional affordances remains a core challenge. We int…
cs.CV2025
HouseLayout3D: A Benchmark and Training-Free Baseline for 3D Layout Estimation in the Wild
Valentin Bieri, Marie-Julie Rakotosaona, Keisuke Tateno +2
Current 3D layout estimation models are primarily trained on synthetic datasets containing simple single room or single floor environments. As a consequence, they cannot natively h…
cs.CV2025
Dynamic Reflections: Probing Video Representations with Text Alignment
Tyler Zhu, Tengda Han, Leonidas Guibas +2
The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encod…