works on

From the 1 of 37 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CLShow all

24 papers · 1 filter

cs.CL2026

Group Entropy-Controlled Policy Optimization

Guangran Cheng, Chengqi Lyu, Songyang Gao +2

Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment pro…

cs.CL2026

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

Lingkai Kong, Zijian Wu, Yuzhe Gu +10

The paper introduces AdvancedMathBench, a benchmark suite for evaluating large language models on generating and verifying advanced mathematical proofs, and provides an automatic v…

cs.CL2026

InternBootcamp: Boosting LLM Reasoning with Verifiable Task Scaling

Peiji Li, Jiasheng Ye, Yongkang Chen +19

Large language models (LLMs) have revolutionized artificial intelligence by enabling complex reasoning capabilities. While recent advancements in reinforcement learning (RL) have p…

cs.CL2026

The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner

Zhouqi Hua, Wenwei Zhang, Chengqi Lyu +5

Length generalization, the ability to solve problems of longer sequences than those observed during training, poses a core challenge of Transformer-based large language models (LLM…

cs.CL2026

Pre-Trained Policy Discriminators are General Reward Models

Shihan Dou, Shichun Liu, Yuming Yang +19

We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guidi…

cs.CL2025

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving

Songyang Gao, Yuzhe Gu, Zijian Wu +18

Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning with Verifiable Rewards (RLVR),…