2 papers
cs.CL2025
Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning
Bowen Jin, TJ Collins, Donghan Yu +10
Large language models (LLMs) exhibit complementary strengths across domains and come with varying inference costs, motivating the design of multi-agent LLM systems where specialize…
cs.LG2025
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
Shenao Zhang, Donghan Yu, Yihao Feng +4
Large language models excel with reinforcement learning (RL), but fully unlocking this potential requires a mid-training stage. An effective mid-training phase should identify a co…