distributional concentration 1language models 1mathematical reasoning 1policy optimization 1reinforcement learning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
ReCo: Reweighting GRPO Against Distributional Concentration
Junoh Park, Junseo Hwang, Wonguk Cho +1
The paper introduces ReCo, a reweighting technique for Group Relative Policy Optimization that mitigates the method’s tendency to focus on high‑probability responses, thereby impro…
cs.LG2026
PiCa: Parameter-Efficient Fine-Tuning with Column Space Projection
Junseo Hwang, Wonguk Cho, Taesup Kim
Fine-tuning large foundation models is essential for building expert models tailored to specialized tasks and domains, but fully updating billions of parameters is computationally…