1 citations · 1 across the 4 of their papers we have counts for
32 papers
Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
Yulin Luo, Chun-Kai Fan, Menghang Dong +19
Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, w…
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
Enshen Zhou, Yibo Li, Jingkun An +12
Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spa…
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
Jiaxin Ai, Tao Hu, Xuemeng Yang +11
Existing computer-use agents remain fundamentally limited in professional software manipulation: GUI-based agents suffer from fragile visual grounding and long-horizon error accumu…
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation
Shuyi Zhang, Yunfan Lou, Hongyang Cheng +8
Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While Reinforcement Learning (RL) fine-tuning can surpass this limit…
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
Cheng Peng, Zhenzhe Zhang, Xiaobao Wei +7
Object navigation in unseen indoor environments requires agents to perform semantic search under partial observability. Vision-language models (VLMs) provide strong semantic-spatia…
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
Lingfeng Zhang, Xiaoshuai Hao, Qinwen Xu +7
Vision-and-language navigation (VLN) is a key task in Embodied AI, requiring agents to navigate diverse and unseen environments while following natural language instructions. Tradi…