1 paper
Pattaraphon Kenny Wongchamcharoen, Kris Gulati, Min Min Fong +1
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives…