Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Rethinking GSPO: The Perplexity-Entropy Equivalence
Chi Liu
We provide a new perspective on GSPO's length-normalized importance ratios by establishing their connection to information-theoretic quantities. We show that GSPO's sequence-level…
cs.LG2025
Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
Chi Liu, Derek Li, Yan Shu +4
While large language models show promise in medical applications, achieving expert-level clinical reasoning remains challenging due to the need for both accurate answers and transp…