3 papers
cs.AI2025
Compass-Thinker-7B Technical Report
Anxiang Zeng, Haibo Zhang, Kaixiang Mo +6
Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning i…
cs.CL2025
dots.llm1 Technical Report
Bi Huo, Bin Tu, Cheng Qin +24
Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this…
cs.AI2025
Establishing Reliability Metrics for Reward Models in Large Language Models
Yizhou Chen, Yawen Liu, Xuesi Wang +5
The reward model (RM) that represents human preferences plays a crucial role in optimizing the outputs of large language models (LLMs), e.g., through reinforcement learning from hu…