1 paper · 1 filter
Arash Afkanpour, Omkar Dige, Fatemeh Tavakoli +3
Current evaluation frameworks for foundation models rely heavily on static, manually curated benchmarks, limiting their ability to capture the full breadth of model capabilities. T…