Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
Yasi Zhang, Tianyu Chen, Mingyuan Zhou +3
Large language models (LLMs) are increasingly deployed as automated evaluators that assign numeric scores to model outputs, a paradigm known as LLM-as-a-Judge. However, standard Re…
cs.LG2025
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
Ling Team, Caizhi Tang, Chilin Fu +15
This technical report presents Ring-Lite-Distill, a lightweight reasoning model derived from our open-source Mixture-of-Experts (MoE) Large Language Models (LLMs) Ling-Lite. This s…