activity
20242026
most citedM2Diffuser: Diffusion-based Trajectory Optimization for Mobile Manipulation in 3D Scenes

14 citations · 14 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2026

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

Jiapeng Li, Ping Wei, Wenjuan Han +2

Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark mat…

cs.RO2026

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation

Zhuofan Zhang, Tianxu Wang, Guoxi Zhang +6

Mobile manipulation requires a robot to navigate to a target object or receptacle and then perform intended manipulation. However, reaching the vicinity of the target does not guar…

cs.RO2025

Path and Motion Optimization for Efficient Multi-Location Inspection with Humanoid Robots

Jiayang Wu, Jiongye Li, Shibowen Zhang +8

This paper proposes a novel framework for humanoid robots to execute inspection tasks with high efficiency and millimeter-level precision. The approach combines hierarchical planni…

cs.RO2025

Integration of Robot and Scene Kinematics for Sequential Mobile Manipulation Planning

Ziyuan Jiao, Yida Niu, Zeyu Zhang +5

We present a Sequential Mobile Manipulation Planning (SMMP) framework that can solve long-horizon multi-step mobile manipulation tasks with coordinated whole-body motion, even when…

cs.RO2024★ 14 cited

M2Diffuser: Diffusion-based Trajectory Optimization for Mobile Manipulation in 3D Scenes

Sixu Yan, Zeyu Zhang, Muzhi Han +7

Recent advances in diffusion models have opened new avenues for research into embodied AI agents and robotics. Despite significant achievements in complex robotic locomotion and sk…

cs.RO2024

M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes

Zeyu Zhang, Sixu Yan, Muzhi Han +4

We propose M3Bench, a new benchmark for whole-body motion generation in mobile manipulation tasks. Given a 3D scene context, M3Bench requires an embodied agent to reason about its…