Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
Hejian Sang, Yuanda Xu, Zhengze Zhou +3
Reasoning models often generate far more tokens than a task requires, which raises inference cost and can compound errors. We introduce CRISP (Compressed Reasoning via Iterative Se…
cs.LG2026
Not all tokens are needed(NAT): token efficient reinforcement learning
Hejian Sang, Yuanda Xu, Zhengze Zhou +2
Reinforcement learning (RL) has become a key driver of progress in large language models, but scaling RL to long chain-of-thought (CoT) trajectories is increasingly constrained by…