activity
20242026
most citedSCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

Haosi Mo, Zihao Yan, Ruiqing Zhang +4

Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external search and browsing tools.…

cs.CL2026

MemoNoveltyAgent: A Historical Research Memory-Aware Agent Workflow for Paper Novelty Assessment

Jiajun Hou, Hexuan Deng, Wenxiang Jiao +4

To alleviate the heavy burden of paper screening, researchers increasingly rely on existing AI agents, such as AI reviewers or DeepResearch, for paper evaluation and novelty assess…

cs.CL2026

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models

Shuo Nie, Hexuan Deng, Chao Wang +6

As large language models become smaller and more efficient, small reasoning models (SRMs) are crucial for enabling chain-of-thought (CoT) reasoning in resource-constrained settings…

cs.CL2026

CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

Hexuan Deng, Xiaopeng Ke, Yichen Li +6

Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews…

cs.CL2026

RouterKGQA: Specialized--General Model Routing for Constraint-Aware Knowledge Graph Question Answering

Bo Yuan, Hexuan Deng, Xuebo Liu +1

Knowledge graph question answering (KGQA) is a promising approach for mitigating LLM hallucination by grounding reasoning in structured and verifiable knowledge graphs. Existing ap…

cs.CL2026

REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning

Hexuan Deng, Wenxiang Jiao, Xuebo Liu +2

Large Reasoning Models (LRMs) demonstrate strong performance in complex tasks but often face the challenge of overthinking, leading to substantially high inference costs. Existing…