activity
20242026
most citedRM-R1: Reward Modeling as Reasoning

2 citations · 2 across the 2 of their papers we have counts for

collaborators
Showing 2025Show all

15 papers · 1 filter

cs.CL2025

Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis

Shuhaib Mehri, Xiusi Chen, Heng Ji +1

High-quality instruction-tuning data is crucial for developing Large Language Models (LLMs) that can effectively navigate real-world tasks and follow human instructions. While synt…

cs.LG2025

SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals

Peixuan Han, Cheng Qian, Xiusi Chen +3

Large language models (LLMs) exhibit exceptional capabilities across various tasks but also pose risks by generating harmful content. Existing safety mechanisms, while improving mo…

cs.CL2025

DecisionFlow: Advancing Large Language Model as Principled Decision Maker

Xiusi Chen, Shanyong Wang, Cheng Qian +3

In high-stakes domains such as healthcare and finance, effective decision-making demands not just accurate outcomes but transparent and explainable reasoning. However, current lang…

cs.AI2025

Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search

Jonathan Light, Min Cai, Weiqin Chen +5

Traditional reinforcement learning and planning typically requires vast amounts of data and training to develop effective policies. In contrast, large language models (LLMs) exhibi…

cs.AI2025

Acting Less is Reasoning More! Teaching Model to Act Efficiently

Hongru Wang, Cheng Qian, Wanjun Zhong +7

Tool-integrated reasoning (TIR) augments large language models (LLMs) with the ability to invoke external tools during long-form reasoning, such as search engines and code interpre…

cs.AI2025

SMART: Self-Aware Agent for Tool Overuse Mitigation

Cheng Qian, Emre Can Acikgoz, Hongru Wang +5

Current Large Language Model (LLM) agents demonstrate strong reasoning and tool use capabilities, but often lack self-awareness, failing to balance these approaches effectively. Th…