#embodied ai

topicembodied ai

18 papers · 1 filter

cs.CV2026

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Yukang Cao, Haozhe Xie, Beichen Wen +13

The paper presents ACE, a data collection system that records synchronized multimodal streams—including egocentric and multi-view video, full-body and hand motion, object geometry,…

cs.CV2026

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Zexuan Yan, Yuzhou Wu, Yue Ma +9

The paper introduces EgoGenesis, a simulator that generates controllable egocentric manipulation videos using geometry-aware conditioning mechanisms to augment real robot data and…

cs.CR2026

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

Fazhong Liu, Zhuoyan Chen, Haozhen Tan +3

The paper surveys security risks and defenses for world-model-based embodied AI, examining how attacks can affect data, perception, prediction, and action throughout the system’s l…

cs.AI2026

The Geometry of Flow-Matching Uncertainty: A Cost-free Uncertainty Proxy and Its Application in Flow-based VLA Failure Detection

Ziyang Rao, Yiren Zhao, Weiyu Guo +3

The paper interprets uncertainty in flow‑matching based action models as geometric deviation in the velocity field and proposes a cost‑free proxy called denoising acceleration that…

cs.AI2026

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Yang Zhou, Zixuan Huang, Sunzhu Li +10

The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…

cs.CV2026

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Hengyi Xie, Chenfei Yao, Xianjin Wu +7

TurboVLA is a vision-language-action model that directly maps visual observations and language instructions to robot actions, achieving real-time performance (32 Hz) on an RTX 4090…