activity
20242026
most citedDo NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

2 citations · 4 across the 19 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning

Haolin Liu, Dian Yu, Sidi Lu +6

Reinforcement learning (RL) has emerged as a powerful framework for improving the reasoning capabilities of large language models (LLMs). However, most existing RL approaches rely…

cs.LG2025

Stable and Efficient Single-Rollout RL for Multimodal Reasoning

Rui Liu, Dian Yu, Lei Ke +6

Reinforcement Learning with Verifiable Rewards (RLVR) has become a key paradigm to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevalen…

cs.LG2025

Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values

Dian Yu, Yulai Zhao, Kishan Panaganti +3

We propose Reinforcement Learning with Explicit Human Values (RLEV), a method that aligns Large Language Model (LLM) optimization directly with quantifiable human value signals. Wh…

cs.LG2025

Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation

Yujun Zhou, Zhenwen Liang, Haolin Liu +7

Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR), yet real-world deployment demands models that can self-improve wit…

cs.LG20251 cited

One Token to Fool LLM-as-a-Judge

Yulai Zhao, Haolin Liu, Dian Yu +4

Large language models (LLMs) are increasingly trusted as automated judges, assisting evaluation and providing reward signals for training other models, particularly in reference-ba…

cs.LG2025

Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Yuheng Zhang, Dian Yu, Tao Ge +5

Reinforcement learning from human feedback (RLHF) has demonstrated remarkable effectiveness in aligning large language models (LLMs) with human preferences. Many existing alignment…