collaborators

6 papers

cs.AI2025

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Kaiyi Zhang, Ang Lv, Jinpeng Li +4

Reinforcement learning with verifiable rewards (RLVR) is a promising approach for improving the complex reasoning abilities of large language models (LLMs). However, current RLVR m…

cs.LG2025

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games

Xiaoqing Zhang, Huabin Zheng, Ang Lv +5

Large language models (LLMs) have been observed to suddenly exhibit advanced reasoning abilities during reinforcement learning (RL), resembling an ``aha moment'' triggered by simpl…

cs.CL2025

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Ang Lv, Ruobing Xie, Xingwu Sun +2

Recent studies on post-training large language models (LLMs) for reasoning through reinforcement learning (RL) typically focus on tasks that can be accurately verified and rewarded…

cs.CL2025

Autonomy-of-Experts Models

Ang Lv, Ruobing Xie, Yining Qian +5

Mixture-of-Experts (MoE) models mostly use a router to assign tokens to specific expert modules, activating only partial parameters and often outperforming dense models. We argue t…

cs.LG2025

More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives

Xiaoqing Zhang, Ang Lv, Yuhan Liu +6

Large language models (LLMs) excel at few-shot in-context learning (ICL) without requiring parameter updates. However, as ICL demonstrations increase from a few to many, performanc…

cs.CL2024

PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead

Tao Tan, Yining Qian, Ang Lv +7

Large language models (LLMs) enhanced with retrieval-augmented generation (RAG) have introduced a new paradigm for web search. However, the limited context awareness of LLMs degrad…