most citedSCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

1 citations · 1 across the 6 of their papers we have counts for

collaborators

12 papers

cs.CL2026

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

Haosi Mo, Zihao Yan, Ruiqing Zhang +4

Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external search and browsing tools.…

cs.LG2026

Distilled Reinforcement Learning for LLM Post-training

Chen Wang, Zhaochun Li, Jionghao Bai +4

Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL)…

cs.CL2026

MemoNoveltyAgent: A Historical Research Memory-Aware Agent Workflow for Paper Novelty Assessment

Jiajun Hou, Hexuan Deng, Wenxiang Jiao +4

To alleviate the heavy burden of paper screening, researchers increasingly rely on existing AI agents, such as AI reviewers or DeepResearch, for paper evaluation and novelty assess…

cs.CL2026

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models

Shuo Nie, Hexuan Deng, Chao Wang +6

As large language models become smaller and more efficient, small reasoning models (SRMs) are crucial for enabling chain-of-thought (CoT) reasoning in resource-constrained settings…

cs.LG20261 cited

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Chen Wang, Zhaochun Li, Jionghao Bai +3

Reinforcement learning (RL) is a key paradigm for post-training large language models (LLMs), but the widely used Group Relative Policy Optimization (GRPO) often suffers from entro…

cs.CL2026

CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

Hexuan Deng, Xiaopeng Ke, Yichen Li +6

Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews…