Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
MemGym: a Long-Horizon Memory Environment for LLM Agents
Wujiang Xu, Yu Wang, Kai Mei +8
Memory is a central capability for LLM agents operating across long-horizon tasks. Existing memory benchmarks predominantly evaluate retention of personalized information in multi-…
cs.CL2025
TeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning
Shangding Gu, Alois Knoll, Ming Jin
The development of Large Language Models (LLMs) often confronts challenges stemming from the heavy reliance on human annotators in the reinforcement learning with human feedback (R…