activity
20242026
collaborators

12 papers

cs.AI2026

Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context Length

Jingxuan Chen, Mohammad Taher Pilehvar, Jose Camacho-Collados

Users often rely on Large Language Models (LLMs) for processing multiple documents or performing analysis over a number of instances. For example, analysing the overall sentiment o…

cs.CL2026

Synthia: Scalable Grounded Persona Generation from Social Media Data

Vahid Rahimzadeh, Erfan Moosavi Monazzah, Mohammad Taher Pilehvar +1

Persona-driven simulations are increasingly used in computational social science, yet their validity critically depends on the fidelity of the underlying personas. Constructing vir…

cs.CL2025

Exploring State Tracking Capabilities of Large Language Models

Kiamehr Rezaee, Jose Camacho-Collados, Mohammad Taher Pilehvar

Large Language Models (LLMs) have demonstrated impressive capabilities in solving complex tasks, including those requiring a certain level of reasoning. In this paper, we focus on…

cs.CL2025

Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets

Mahdi Zakizadeh, Mohammad Taher Pilehvar

Accurately measuring gender stereotypical bias in language models is a complex task with many hidden aspects. Current benchmarks have underestimated this multifaceted challenge and…

cs.CL2025

Pun Unintended: LLMs and the Illusion of Humor Understanding

Alessandro Zangari, Matteo Marcuzzo, Andrea Albarelli +2

Puns are a form of humorous wordplay that exploits polysemy and phonetic similarity. While LLMs have shown promise in detecting puns, we show in this paper that their understanding…

cs.CL2025

MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables

Matteo Marcuzzo, Alessandro Zangari, Andrea Albarelli +2

As LLMs excel on standard reading comprehension benchmarks, attention is shifting toward evaluating their capacity for complex abstract reasoning and inference. Literature-based be…