1 paper · 1 filter
Ritvik Garimella, Vedant Khandelwal, Anvi Kohli +1
Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema w…