#embodied AI

try —

18 papers match

cs.CR2026

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

Fazhong Liu, Zhuoyan Chen, Haozhen Tan +3

The paper surveys security risks and defenses for world-model-based embodied AI, examining how attacks can affect data, perception, prediction, and action throughout the system’s l…

#embodied ai#world models#adversarial attacks#defense strategies
cs.AI2026

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Yang Zhou, Zixuan Huang, Sunzhu Li +10

The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…

#vision-language models#spatial reasoning#tool use#embodied AI
cs.AI2026

The Geometry of Flow-Matching Uncertainty: A Cost-free Uncertainty Proxy and Its Application in Flow-based VLA Failure Detection

Ziyang Rao, Yiren Zhao, Weiyu Guo +3

The paper interprets uncertainty in flow‑matching based action models as geometric deviation in the velocity field and proposes a cost‑free proxy called denoising acceleration that…

#flow matching#uncertainty estimation#embodied AI#online failure detection
cs.CV2026

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Zexuan Yan, Yuzhou Wu, Yue Ma +9

The paper introduces EgoGenesis, a simulator that generates controllable egocentric manipulation videos using geometry-aware conditioning mechanisms to augment real robot data and…

#egocentric video#embodied AI#video synthesis#robot manipulation
cs.CV2026

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Yukang Cao, Haozhe Xie, Beichen Wen +13

The paper presents ACE, a data collection system that records synchronized multimodal streams—including egocentric and multi-view video, full-body and hand motion, object geometry,…

#embodied ai#multimodal dataset#human-centric capture#egocentric video
cs.CV2026

HumanCLAW: Can Vision-Language Models Act Through a Body?

Siyao Li, Li Siyao, Jiawei Gu +16

The paper introduces HumanCLAW, a framework that separates decision making of vision‑language models from low‑level motor execution, allowing evaluation of a model's action intelli…

#vision-language models#embodied AI#benchmarking#physical simulation