Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
INTELLECT-3: Technical Report
Prime Intellect Team, Mika Senghaas, Fares Obeid +20
We present INTELLECT-3, a 106B-parameter Mixture-of-Experts model (12B active) trained with large-scale reinforcement learning on our end-to-end RL infrastructure stack. INTELLECT-…
cs.LG2025
Reinforcing Multi-Turn Reasoning in LLM Agents via Fine-Grained Reward Structure and Credit Assignment
Quan Wei, Siliang Zeng, Chenliang Li +9
Reinforcement Learning (RL) approaches have been wildly used to enhance the reasoning capabilities of Large Language Model (LLM) agents in long-horizon, multi-turn scenarios. Such…
cs.LG2024
Online Stackelberg Optimization via Nonlinear Control
William Brown, Christos Papadimitriou, Tim Roughgarden
In repeated interaction problems with adaptive agents, our objective often requires anticipating and optimizing over the space of possible agent responses. We show that many proble…