9 papers
DemoBridge: A Simulation-in-the-Loop Toolkit for Single-View Human Demonstration Retargeting
Zehao Wang, Fabien Despinoy, Sergey Zakharov +2
We present DemoBridge, an toolkit that turns a single-view RGB stereo recording of a human hand demonstration into an executable, physics-validated robot-arm trajectory. Retargetin…
Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search
Imen Mahdi, Matteo Cassinelli, Fabien Despinoy +2
Open-world interactive object search in household environments requires understanding semantic relationships between objects and their surrounding context to guide exploration effi…
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
Christian Gumbsch, Leonardo Barcellona, Lennard Schünemann +7
Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics. Recent work has explored…
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
Andrii Zadaianchuk, Leonardo Barcellona, Lennard Schuenemann +7
Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulati…
Object Pose Transformer: Unifying Unseen Object Pose Estimation
Weihang Li, Lorenzo Garattoni, Fabien Despinoy +2
Learning model-free object pose estimation for unseen instances remains a fundamental challenge in 3D vision. Existing methods typically fall into two disjoint paradigms: category-…
Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge
Li Kang, Heng Zhou, Xiufeng Song +41
Recent advancements in multimodal large language models and vision-languageaction models have significantly driven progress in Embodied AI. As the field transitions toward more com…