5 papers
RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs
Runlong Zhou, Lefan Zhang, Shang-Chen Wu +29
Reinforcement learning (RL) has emerged as the de-facto paradigm for improving the reasoning capabilities of large language models (LLMs). We have developed RLAX, a scalable RL fra…
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
Shenao Zhang, Donghan Yu, Yihao Feng +4
Large language models excel with reinforcement learning (RL), but fully unlocking this potential requires a mid-training stage. An effective mid-training phase should identify a co…
Text2Data: Low-Resource Data Generation with Textual Control
Shiyu Wang, Yihao Feng, Tian Lan +6
Natural language serves as a common and straightforward signal for humans to interact seamlessly with machines. Recognizing the importance of this interface, the machine learning c…
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
Haolin Chen, Yihao Feng, Zuxin Liu +8
Large language models (LLMs) have shown impressive capabilities, but still struggle with complex reasoning tasks requiring multiple steps. While prompt-based methods like Chain-of-…
AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning
Jianguo Zhang, Tian Lan, Rithesh Murthy +15
Autonomous agents powered by large language models (LLMs) have garnered significant research attention. However, fully harnessing the potential of LLMs for agent-based tasks presen…