1 paper
Serina Chang, Ashton Anderson, Jake M. Hofman
With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together. However, standard benchmarks, such as MMLU, measure L…