#embodied ai
18 papers · 1 filter
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Yukang Cao, Haozhe Xie, Beichen Wen +13
The paper presents ACE, a data collection system that records synchronized multimodal streams—including egocentric and multi-view video, full-body and hand motion, object geometry,…
EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
Zexuan Yan, Yuzhou Wu, Yue Ma +9
The paper introduces EgoGenesis, a simulator that generates controllable egocentric manipulation videos using geometry-aware conditioning mechanisms to augment real robot data and…
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
Fazhong Liu, Zhuoyan Chen, Haozhen Tan +3
The paper surveys security risks and defenses for world-model-based embodied AI, examining how attacks can affect data, perception, prediction, and action throughout the system’s l…
The Geometry of Flow-Matching Uncertainty: A Cost-free Uncertainty Proxy and Its Application in Flow-based VLA Failure Detection
Ziyang Rao, Yiren Zhao, Weiyu Guo +3
The paper interprets uncertainty in flow‑matching based action models as geometric deviation in the velocity field and proposes a cost‑free proxy called denoising acceleration that…
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Yang Zhou, Zixuan Huang, Sunzhu Li +10
The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Hengyi Xie, Chenfei Yao, Xianjin Wu +7
TurboVLA is a vision-language-action model that directly maps visual observations and language instructions to robot actions, achieving real-time performance (32 Hz) on an RTX 4090…