3 papers
cs.CL2026
LLMs Can Better Capture Human Judgments--With the Right Prompts
Danica Dillion, Chen Cecilia Liu, Baihui Wang +5
Are large language models (LLMs) bad at capturing human judgment? Two commonly stated limitations are that LLMs fail to capture full distributions of responses, and that their judg…
cs.CL2026
From Tokens To Agents: A Researcher's Guide To Understanding Large Language Models
Daniele Barolo
Researchers face a critical choice: how to use -- or not use -- large language models in their work. Using them well requires understanding the mechanisms that shape what LLMs can…
cs.CY2025
Whose Name Comes Up? Auditing LLM-Based Scholar Recommendations
Daniele Barolo, Chiara Valentin, Fariba Karimi +3
This paper evaluates the performance of six open-weight LLMs (llama3-8b, llama3.1-8b, gemma2-9b, mixtral-8x7b, llama3-70b, llama3.1-70b) in recommending experts in physics across f…