4 papers
ASH: Agents that Self-Hone via Embodied Learning
Benjamin Schneider, Xavier Schneider, Victor Zhong +1
Long-horizon embodied tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, neither of which scales. We i…
Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies
Sitao Cheng, Xunjian Yin, Ruiwen Zhou +5
Does Reinforcement Learning (RL) merely amplify existing skills, or synthesize novel skills? We investigate this question through the lens of Complementary Reasoning: the critical…
How well can LLMs provide planning feedback in grounded environments?
Yuxuan Li, Victor Zhong
Learning to plan in grounded environments typically requires carefully designed reward functions or high-quality annotated demonstrations. Recent works show that pretrained foundat…
Policy Improvement using Language Feedback Models
Victor Zhong, Dipendra Misra, Xingdi Yuan +1
We introduce Language Feedback Models (LFMs) that identify desirable behaviour - actions that help achieve tasks specified in the instruction - for imitation learning in instructio…