Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
GIM: Evaluating models via tasks that integrate multiple cognitive domains
Rohit Patel, Alexandre Rezende, Steven McClain
As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or removing knowledge entirely in f…
cs.AI2024
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +556
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models th…