Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Xiaoying Zhang, Yipeng Zhang, Hao Sun +4
Recent advances in reinforcement learning (RL) using numerical rewards have significantly enhanced the complex reasoning capabilities of large language models (LLMs). However, we i…
cs.CL2025
The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data
Thomas Pouplin, Katarzyna Kobalczyk, Hao Sun +1
Developing autonomous agents capable of performing complex, multi-step decision-making tasks specified in natural language remains a significant challenge, particularly in realisti…
cs.CL2024
Retrieval Augmented Thought Process for Private Data Handling in Healthcare
Thomas Pouplin, Hao Sun, Samuel Holt +1
Large Language Models (LLMs) have demonstrated the strong potential to assist both clinicians and the general public with their extensive medical knowledge. However, their applicat…