Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
Zeyuan Wang, Da Li, Yulin Chen +4
We introduce a one-step generative policy for offline reinforcement learning that maps noise directly to actions via a residual reformulation of MeanFlow, making it compatible with…
cs.LG2025
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
Yufei He, Ruoyu Li, Alex Chen +8
Large language model (LLM) agents often struggle in environments where rules and required domain knowledge frequently change, such as regulatory compliance and user risk screening.…