1 paper
Fengyuan Liu, Jay Gala, Nilaksh +3
Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on d…