#robot learning

try —

7 papers match

cs.CV2026

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

Weiquan Lin, Yu Deng, Shiyang Liu +6

This survey systematically reviews how large foundation models provide geometric, semantic, and visual priors for hand‑object interaction tasks such as reconstruction and generatio…

#hand-object interaction#foundation models#reconstruction#generation
cs.AI2026

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

Jia Luo

The paper presents Pegasus, a framework that converts human manipulation videos into robot-learnable data by building intermediate graph representations and verifying them with phy…

#embodied ai#robot learning#video-to-robot translation#affordance modeling
cs.RO2026

VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents

Marcus Hoerger, Rishikesh Joshi, Rahul Shome +2

VOiLA learns task‑agnostic POMDP transition and observation models with conditional diffusion networks, distills them into fast feed‑forward generators, and integrates them with a…

#pomdp planning#diffusion models#online planning#belief updates
cs.RO2026

Semantic Anchoring for Robotic Action Representations

Yuan Xu, Youheng Shi, Chengyang Li +2

The paper studies how fine‑tuning vision‑language‑action models for robots can degrade the semantic structure of their action representations, and proposes a plug‑and‑play anchorin…

#action representation#vision-language models#semantic anchoring#robot learning
cs.CV2026

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

Yitong Chen, Shiduo Zhang, Jingjing Gong +1

The paper proposes a one-step action generation method for vision‑language‑action models, using high‑noise training and a flow‑matching loss, and demonstrates strong performance on…

#vision-language-action#one-step action generation#flow matching#robot learning
cs.RO2026

Towards Predictive, Aligned, and Scalable Robot Learning

Peijun Tang, Shangjin Xie, Baifu Huang +6

The paper introduces Lumo-2, a latent world-action model that reasons about future physical dynamics in a shared latent space to generate robot actions, using a multi‑stage alignme…

#robot learning#multimodal alignment#latent dynamics#predictive reasoning