1 paper · 1 filter
Dae Yon Hwang, Raunaq Suri, Valentin Villecroze +4
LLM agents operate in two distinct regimes: open-weight agents amenable to reinforcement learning (RL) and black-box agents whose behaviour must be controlled purely at test time.…