2 papers
cs.LG2026
Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training
William Hoy, Binxu Wang, Xu Pan
Evolution Strategies (ES) have emerged as a scalable gradient-free alternative to reinforcement learning based LLM fine-tuning, but it remains unclear whether comparable task perfo…
cs.LG2025
STABLE: Gated Continual Learning for Large Language Models
William Hoy, Nurcin Celik
Large language models (LLMs) increasingly require mechanisms for continual adaptation without full retraining. However, sequential updates can lead to catastrophic forgetting, wher…