most citedLearning to Select In-Context Demonstration Preferred by Large Language Model

1 citations · 1 across the 4 of their papers we have counts for

collaborators

8 papers

cs.AI2026

Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient

Changming Li, Kaixing Zhang, Haoyun Xu +4

Large language models (LLMs) demonstrate strong reasoning abilities in solving complex real-world problems. Yet, the internal mechanisms driving these complex reasoning behaviors r…

cs.LG2026

Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning

Zheng Zhang, Ao Lu, Yuanhao Zeng +5

Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant breakthroughs in complex LLM reasoning within verifiable domains, such as mathematics and programmin…

cs.AI2025

SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models

Jiayi Tian, Seyedarmin Azizi, Yequan Zhao +7

Large reasoning models (LRMs) often incur significant key-value (KV) cache overhead, due to their linear growth with the verbose chain-of-thought (CoT) reasoning. This incurs both…

cs.LG2025

Better LLM Reasoning via Dual-Play

Zhengxin Zhang, Chengyu Huang, Aochong Oliver Li +1

Large Language Models (LLMs) have achieved remarkable progress through Reinforcement Learning with Verifiable Rewards (RLVR), yet still rely heavily on external supervision (e.g.,…

cs.LG2025

Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning

Zheng Zhang, Ziwei Shan, Kaitao Song +2

Process Reward Models (PRMs) have emerged as a promising approach to enhance the reasoning capabilities of large language models (LLMs) by guiding their step-by-step reasoning towa…

cs.AI2025

Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning

Zheng Zhang

Large Language Models (LLMs) display striking surface fluency yet systematically fail at tasks requiring symbolic reasoning, arithmetic accuracy, and logical consistency. This pape…