4 papers
Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation
Luca Zhou, Sajel Shah, Emanuele Rodolà +1
Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficulty signal. The same signal drives RL with…
Generative AI collective behavior needs an interactionist paradigm
Laura Ferrarotti, Gian Maria Campedelli, Roberto Dessì +7
In this article, we argue that understanding the collective behavior of agents based on large language models (LLMs) is an essential area of inquiry, with important implications in…
Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
Davide Marincione, Donato Crisostomi, Roberto Dessi +2
Foundation models capable of generalizing across species and tasks represent a promising new frontier in bioacoustics, with NatureLM being one of the most prominent examples. While…
I Want to Break Free! Persuasion and Anti-Social Behavior of LLMs in Multi-Agent Settings with Social Hierarchy
Gian Maria Campedelli, Nicolò Penzo, Massimo Stefan +4
As LLM-based agents become increasingly autonomous and will more freely interact with each other, studying the interplay among them becomes crucial to anticipate emergent phenomena…