4 papers · 1 filter
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
Zhihan Liu, Lin Guan, Yixin Nie +6
Generalist LLM agents are often post-trained on a narrow set of environments but deployed across far broader, unseen domains. In this work, we investigate the challenge of agentic…
Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following
Yuxiao Yang, Shenao Zhang, Zhihan Liu +2
This work focuses on building a task planner for Embodied Instruction Following (EIF) using Large Language Models (LLMs). Previous works typically train a planner to imitate expert…
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
Ruijie Xu, Zhihan Liu, Yongfei Liu +4
We address the challenge of online Reinforcement Learning from Human Feedback (RLHF) with a focus on self-rewarding alignment methods. In online RLHF, obtaining feedback requires i…
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
Zhihan Liu, Hao Hu, Shenao Zhang +4
Large language models (LLMs) demonstrate impressive reasoning abilities, but translating reasoning into actions in the real world remains challenging. In particular, it remains unc…