6 papers
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
Hao Yi, Yulan Hu, Xin Li +3
Large Language Models (LLMs) have recently improved mathematical reasoning through Reinforcement Learning with Verifiable Reward (RLVR). However, existing RLVR algorithms require l…
Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
Ruifeng Ren, Sheng Ouyang, Huayi Tang +1
Attention-based Transformers have demonstrated strong adaptability across a wide range of tasks and have become the backbone of modern Large Language Models (LLMs). However, their…
AMAP Agentic Planning Technical Report
AMAP AI Agent Team, Yulan Hu, Xiangwen Zhang +22
We present STAgent, an agentic large language model tailored for spatio-temporal understanding, designed to solve complex tasks such as constrained point-of-interest discovery and…
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
Yulan Hu, Sheng Ouyang, Jinman Zhao +1
The Process Reward Model (PRM) plays a crucial role in mathematical reasoning tasks, requiring high-quality supervised process data. However, we observe that reasoning steps genera…
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
Sheng Ouyang, Yulan Hu, Ge Chen +3
Rewards serve as proxies for human preferences and play a crucial role in Reinforcement Learning from Human Feedback (RLHF). However, if these rewards are inherently imperfect, exh…
GUNDAM: Aligning Large Language Models with Graph Understanding
Sheng Ouyang, Yulan Hu, Ge Chen +1
Large Language Models (LLMs) have achieved impressive results in processing text data, which has sparked interest in applying these models beyond textual data, such as graphs. In t…