3 papers
cs.LG2026
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18
Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…
cs.RO2026
Swim2Real: VLM-Guided System Identification for Sim-to-Real Transfer
Kevin Qiu, Kyle Walker, Mike Y. Michelis +2
We present Swim2Real, a pipeline that calibrates a 16-parameter robotic fish simulator from swimming videos using vision-language model (VLM) feedback, requiring no hand-designed s…
cs.RO2026
Vid2Sid: Videos Can Help Close the Sim2Real Gap
Kevin Qiu, Yu Zhang, Marek Cygan +1
Calibrating a robot simulator's physics parameters (friction, damping, material stiffness) to match real hardware is often done by hand or with black-box optimizers that reduce err…