Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
Yufei He, Ruoyu Li, Alex Chen +8
Large language model (LLM) agents often struggle in environments where rules and required domain knowledge frequently change, such as regulatory compliance and user risk screening.…
cs.LG2025
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
Simon Sinong Zhan, Qingyuan Wu, Philip Wang +4
Offline-to-online deployment of reinforcement-learning (RL) agents must bridge two gaps: (1) the sim-to-real gap, where real systems add latency and other imperfections not present…
cs.LG2024
Inverse Delayed Reinforcement Learning
Simon Sinong Zhan, Qingyuan Wu, Zhian Ruan +6
Inverse Reinforcement Learning (IRL) has demonstrated effectiveness in a variety of imitation tasks. In this paper, we introduce an IRL framework designed to extract rewarding feat…