activity
20242026
most citedA Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

1 citations · 1 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2026

Fast, Slow, and Tool-augmented Thinking for LLMs: A Review

Xinda Jia, Jinpeng Li, Zezhong Wang +6

Large Language Models (LLMs) have demonstrated remarkable progress in reasoning across diverse domains. However, effective reasoning in real-world tasks requires adapting the reaso…

cs.CL2026

BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation

Ning Li, Zixuan Guo, Yan Xu +7

Hallucinations remain a major obstacle to deploying large language models (LLMs) in knowledge-intensive settings, where generated responses must be faithfully grounded in provided…

cs.CL2026

LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

Aofan Yu, Chenyu Zhou, Tianyi Xu +8

Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substantial context overhead and e…

cs.CL2026

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

Haoyi Hu, Qirong Lyu, Xianghan Kong +7

While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after explicit user prompts. This par…

cs.CL2026

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning

Congmin Zheng, Jiachen Zhu, Jianghao Lin +6

Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. Howeve…

cs.CL2026

Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs

Tara Azin, Yongan Yu, Raj Singh +1

Presupposition projection in conditionals is central to theories of meaning and pragmatics, yet it remains largely unevaluated in large language models. We address this gap through…