5 papers
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis
Yu Qi, Haibo Zhao, Ziyu Guo +17
Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly…
Sampling-Based Multi-Modal Multi-Robot Multi-Goal Path Planning
Valentin N. Hartmann, Tirza Heinle, Yijiang Huang +1
In many robotics applications, multiple robots are working in a shared workspace to complete a set of tasks as fast as possible. Such settings can be treated as multi-modal multi-r…
Teaching Robots Like Dogs: Learning Agile Navigation from Luring, Gesture, and Speech
Taerim Yoon, Dongho Kang, Jin Cheng +5
In this work, we aim to enable legged robots to learn how to interpret human social cues and produce appropriate behaviors through physical human guidance. However, learning throug…
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
Liujian Tang, Shaokang Dong, Yijia Huang +21
This paper presents MagicGUI, a foundational mobile GUI agent designed to address critical challenges in perception, grounding, and reasoning within real-world mobile GUI environme…
Budget-optimal multi-robot layout design for box sorting
Peiyu Zeng, Yijiang Huang, Simon Huber +1
Robotic systems are routinely used in the logistics industry to enhance operational efficiency, but the design of robot workspaces remains a complex and manual task, which limits t…