1 paper · 1 filter
David Gringras, Misha Salahshoor
Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do. That literature answers a related, but consequentially different, question: what…