1 citations · 2 across the 20 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
MemoryWalker: Stop Training Agents on Contexts They Never Saw
Zinco J, Xunjie Zhu, Shen Huang +3
Production agent harnesses such as Claude Code and Qwen-Agent compress context during rollout, but training under compression creates a conditioning problem: every eviction branche…
cs.LG2026
ESPO: Early-Stopping Proximal Policy Optimization
Zihang Li, Rui Zhou, Yingcheng Shi +8
When a large language model under reinforcement learning commits a wrong reasoning step early in a trajectory, standard algorithms force it to keep generating until the maximum hor…
cs.LG2026
FGGM: Fisher-Guided Gradient Masking for Continual Learning
Chao-Hong Tan, Qian Chen, Wen Wang +6
Catastrophic forgetting impairs the continuous learning of large language models. We propose Fisher-Guided Gradient Masking (FGGM), a framework that mitigates this by strategically…