activity
20242026
most citedFree Process Rewards without Process Labels

1 citations · 2 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG2026

LAD: Learning Advantage Distribution for Reasoning

Wendi Li, Sharon Li

Current reinforcement learning objectives for large-model reasoning primarily focus on maximizing expected rewards. This paradigm can lead to overfitting to dominant reward signals…

cs.AI2026

Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities

Changdae Oh, Seongheon Park, To Eun Kim +8

Uncertainty quantification (UQ) for large language models (LLMs) is a key building block for safety guardrails of daily LLM applications. Yet, even as LLM agents are increasingly d…

cs.LG2025

General Exploratory Bonus for Optimistic Exploration in RLHF

Wendi Li, Changdae Oh, Sharon Li

Optimistic exploration is central to improving sample efficiency in reinforcement learning with human feedback, yet existing exploratory bonus methods to incentivize exploration of…

cs.LG2025

Process Reinforcement through Implicit Rewards

Ganqu Cui, Lifan Yuan, Zefan Wang +22

Dense process rewards have proven a more effective alternative to the sparse outcome-level rewards in the inference-time scaling of large language models (LLMs), particularly in ta…

cs.LG20241 cited

Free Process Rewards without Process Labels

Lifan Yuan, Wendi Li, Huayu Chen +6

Different from its counterpart outcome reward models (ORMs), which evaluate the entire responses, a process reward model (PRM) scores a reasoning trajectory step by step, providing…

cs.CR20241 cited

FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks

Jiongxiao Wang, Fangzhou Wu, Wendi Li +5

Large language models (LLMs) have been widely deployed as the backbone with additional tools and text information for real-world applications. However, integrating external informa…