2 papers
cs.CY2025
Assessing the Reliability and Validity of Large Language Models for Automated Assessment of Student Essays in Higher Education
Andrea Gaggioli, Giuseppe Casaburi, Leonardo Ercolani +3
This study investigates the reliability and validity of five advanced Large Language Models (LLMs), Claude 3.5, DeepSeek v2, Gemini 2.5, GPT-4, and Mistral 24B, for automated essay…
cs.AI2025
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
Fabrizio Davide, Pietro Torre, Leonardo Ercolani +1
We tasked 16 state-of-the-art large language models (LLMs) with estimating the likelihood of Artificial General Intelligence (AGI) emerging by 2030. To assess the quality of these…