Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
Yuchen Li, Cong Lin, Muhammad Umair Nasir +3
We introduce GVGAI-LLM, a video game benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). Built on the General Video Game AI fra…
cs.AI2024
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +556
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models th…