3 papers
cs.AI2026
Context-Picker: Dynamic context selection using multi-stage reinforcement learning
Siyuan Zhu, Chengdong Xu, Kaiqiang Ke +1
In long-context question answering, selecting the appropriate scope of context for a query remains a key and unresolved challenge. Insufficient context can lead to missing essentia…
cs.AI2025
HR: Hierarchical Hindsight Reflection for Multi-Task LLM Agents
Shicheng Ye, Chao Yu, Kaiqiang Ke +2
Large language model (LLM)-based agents have shown strong potential in multi-task scenarios, owing to their ability to transfer knowledge across diverse tasks. However, existing ap…
cs.LG2025
GCHR : Goal-Conditioned Hindsight Regularization for Sample-Efficient Reinforcement Learning
Xing Lei, Wenyan Yang, Kaiqiang Ke +4
Goal-conditioned reinforcement learning (GCRL) with sparse rewards remains a fundamental challenge in reinforcement learning. While hindsight experience replay (HER) has shown prom…