10 citations · 42 across the 19 of their papers we have counts for
4 papers · 1 filter
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization
Jiacai Liu, Chaojie Wang, Chris Yuhao Liu +5
The role of reinforcement learning (RL) in enhancing the reasoning of large language models (LLMs) is becoming increasingly significant. Despite the success of RL in many scenarios…
Mars-PO: Multi-Agent Reasoning System Preference Optimization
Xiaoxuan Lou, Chaojie Wang, Bo An
Mathematical reasoning is a fundamental capability for large language models (LLMs), yet achieving high performance in this domain remains a significant challenge. The auto-regress…
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Chris Yuhao Liu, Liang Zeng, Jiacai Liu +6
In this report, we introduce a collection of methods to enhance reward modeling for LLMs, focusing specifically on data-centric techniques. We propose effective data selection and…
Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning
Chaojie Wang, Yanchen Deng, Zhiyi Lyu +4
Large Language Models (LLMs) have demonstrated impressive capability in many natural language tasks. However, the auto-regressive generation process makes LLMs prone to produce err…