4 citations · 8 across the 12 of their papers we have counts for
Showing 2025 · cs.CLShow all
2 papers · 2 filters
cs.CL2025
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It
Seyed Mahed Mousavi, Edoardo Cecchinato, Lucia Hornikova +1
We conduct a systematic audit of three widely used reasoning benchmarks, SocialIQa, FauxPas-EAI, and ToMi, and uncover pervasive flaws in both benchmark items and evaluation method…
cs.CL2025
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
Seyed Mahed Mousavi, Simone Alghisi, Giuseppe Riccardi
LLMs' sources of knowledge are data snapshots containing factual information about entities collected at different timestamps and from different media types (e.g. wikis, social med…