1 paper
John Mavi, Nathan Summers, Sergio Coronado
The current paper presents the development and validation of SelfScore, a novel benchmark designed to assess the performance of automated Large Language Model (LLM) agents on help…