activity
20232026
most citedInternLM2 Technical Report

29 citations · 70 across the 41 of their papers we have counts for

collaborators
Showing cs.CLShow all

25 papers · 1 filter

cs.CL2026

Group Entropy-Controlled Policy Optimization

Guangran Cheng, Chengqi Lyu, Songyang Gao +2

Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment pro…

cs.CL2026

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

Lingkai Kong, Zijian Wu, Yuzhe Gu +10

Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly unders…

cs.CL2025

OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification

Zijian Wu, Lingkai Kong, Wenwei Zhang +12

Large language models (LLMs) have achieved significant progress in solving complex reasoning tasks by Reinforcement Learning with Verifiable Rewards (RLVR). This advancement is als…

cs.CL2025

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving

Yuzhe Gu, Songyang Gao, Zijian Wu +18

Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning with Verifiable Rewards (RLVR),…

cs.CL2025

CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward

Shudong Liu, Hongwei Liu, Junnan Liu +8

Answer verification is crucial not only for evaluating large language models (LLMs) by matching their unstructured outputs against standard answers, but also serves as the reward m…

cs.CL2025

InternBootcamp: Boosting LLM Reasoning with Verifiable Task Scaling

Peiji Li, Jiasheng Ye, Yongkang Chen +12

Large language models (LLMs) have revolutionized artificial intelligence by enabling complex reasoning capabilities. While recent advancements in reinforcement learning (RL) have p…