activity
20242026
most citedRobobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

1 citations · 1 across the 21 of their papers we have counts for

collaborators
Showing cs.ROShow all

30 papers · 1 filter

cs.RO2026

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

Jingkai Wang, Zihan Tang, Gu Zhang +7

Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this c…

cs.RO20261 cited

Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

Yulin Luo, Chun-Kai Fan, Menghang Dong +19

Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, w…

cs.RO2026

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

Enshen Zhou, Yibo Li, Jingkun An +12

Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spa…

cs.RO2026

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation

Shuyi Zhang, Yunfan Lou, Hongyang Cheng +8

Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While Reinforcement Learning (RL) fine-tuning can surpass this limit…

cs.RO2026

Dexora: Open-source VLA for High-DoF Bimanual Dexterity

Zongzheng Zhang, Jingrui Pang, Zhuo Yang +22

Vision-Language-Action (VLA) models have recently become a central direction in embodied AI, but current systems are restricted to either dual-gripper control or single-arm dextero…

cs.RO2026

MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation

Lingfeng Zhang, Xiaoshuai Hao, Qinwen Xu +7

Vision-and-language navigation (VLN) is a key task in Embodied AI, requiring agents to navigate diverse and unseen environments while following natural language instructions. Tradi…