3 papers
cs.RO2026
R2S-EGO: Dual-Proxy Refinement for Sparse-Capture Real-to-Sim
Shuai Fang, Xin Deng, Yuchen Kang +2
Real-to-sim (R2S) depends on scene representations that render observations along robot ego trajectories, yet dense multi-view capture limits per-environment real-image capture-cou…
cs.AI2026
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Ying Chen, Weizhen Li, Zhe Hu +7
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state…
cs.RO2026
Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline
Qing Yang, Xun Wang, Ziguan Wang +3
Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed…