8 citations · 8 across the 1 of their papers we have counts for
1 paper
Alex Tamkin, Kunal Handa, Avash Shrestha +1
Language models have recently achieved strong performance across a wide range of NLP benchmarks. However, unlike benchmarks, real world tasks are often poorly specified, and agents…