2 papers
cs.LG2026
Minimal-Intervention KV Retention via Set-Conditioned Diversity
Libo Sun, Po-wei Harn, Peixiong He +1
KV-cache compression at small budgets is a crowded design space spanning cache representation, head-wise routing, compression cadence, decoding behavior, and within-budget scoring.…
cs.AI2025
Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
Jan Niklas Groeneveld, Xi Qin, Alexander Schaefer +1
Generating high-quality code remains a challenge for Large Language Models (LLMs). For the evolution of reasoning models on this task, reward models are a necessary intermediate st…