6 papers
Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents
Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt +5
Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon e…
DuoBench: A Reproducible Benchmark for Bimanual Manipulation in Simulation and the Real World
Tobias Jülg, Seongjin Bien, Simon Hilber +7
Bimanual robot systems substantially expand manipulation capabilities, but coordinating two arms introduces additional control complexity and failure modes that are not well captur…
iPack: Intuitive Bin Packing with Large Language Models
Yannik Blei, Michael Krawez, Adrian Göà +5
Robotics and automation are increasingly influential in logistics but remain largely confined to traditional warehouses. In grocery retail, advancements such as cashier-less superm…
Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale
Tobias Jülg, Pierre Krack, Seongjin Bien +7
Vision-Language-Action models (VLAs) mark a major shift in robot learning. They replace specialized architectures and task-tailored components of expert policies with large-scale d…
Lan-grasp: Using Large Language Models for Semantic Object Grasping and Placement
Reihaneh Mirjalili, Michael Krawez, Yannik Blei +3
In this paper, we propose Lan-grasp, a novel approach towards more appropriate semantic grasping and placing. We leverage foundation models to equip the robot with a semantic under…
CloudTrack: Scalable UAV Tracking with Cloud Semantics
Yannik Blei, Michael Krawez, Nisarga Nilavadi +2
Nowadays, unmanned aerial vehicles (UAVs) are commonly used in search and rescue scenarios to gather information in the search area. The automatic identification of the person sear…