16 papers
Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents
Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt +5
Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon e…
DuoBench: A Reproducible Benchmark for Bimanual Manipulation in Simulation and the Real World
Tobias Jülg, Seongjin Bien, Simon Hilber +7
Bimanual robot systems substantially expand manipulation capabilities, but coordinating two arms introduces additional control complexity and failure modes that are not well captur…
iPack: Intuitive Bin Packing with Large Language Models
Yannik Blei, Michael Krawez, Adrian Göà +5
Robotics and automation are increasingly influential in logistics but remain largely confined to traditional warehouses. In grocery retail, advancements such as cashier-less superm…
Shape Formation for the Cooperative Transportation of Arbitrary Objects Using Multi-Agent Reinforcement Learning
Mohamed Sayed, Wolfram Burgard, Tanja Katharina Kaiser
Cooperative object transportation is essential in numerous domains, including industrial to domestic services. A popular transportation strategy is to carry objects on top of multi…
OGScene3D: Incremental Open-Vocabulary 3D Gaussian Scene Graph Mapping for Scene Understanding
Siting Zhu, Ziyun Lu, Guangming Wang +5
Open-vocabulary scene understanding is crucial for robotic applications, enabling robots to comprehend complex 3D environmental contexts and supporting various downstream tasks suc…
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
Pierre Krack, Tobias Jülg, Wolfram Burgard +1
Well-designed dense reward functions in robot manipulation not only indicate whether a task is completed but also encode progress along the way. Generally, designing dense rewards…