Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Shorter but not Worse: Frugal Reasoning via Easy Samples as Length Regularizers in Math RLVR
Abdelaziz Bounhar, Hadi Abdine, Evan Dufraisse +5
Large language models (LLMs) trained for step-by-step reasoning often become excessively verbose, raising inference cost. Standard Reinforcement Learning with Verifiable Rewards (R…
cs.LG2025
MixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts Models
Ahmad Chamma, Omar El Herraoui, Guokan Shang
We introduce MixtureKit, a modular open-source framework for constructing, training, and analyzing Mixture-of-Experts (MoE) models from arbitrary pre-trained or fine-tuned models.…