2 papers
cs.AI2026
The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
Emma Yanyang Kong, JJ Tan, Ishan Gupta +8
LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accel…
cs.LG2026
Demystifying Reinforcement Learning Post-Training of Language Models
Donovan Clay, Saket Gollapudi, Sankar Harilal +4
Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, a…