1 paper
Stefan Broecker, Mason del Rosario, Boris Selitser +1
The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in the training and evaluation of these mod…