3 papers
cs.CY2026
Efficient Safety Benchmarking via Item Response Theory
Fabio Spagliardi, MÃrian Silva, Ayan Datta +3
Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly…
cs.CL2026
Large Language Models Decide Early and Explain Later
Ayan Datta, Zhixue Zhao, Bhuvanesh Verma +3
Large Language Models often achieve strong performance by generating long intermediate chain-of-thought reasoning. However, it remains unclear when a model's final answer is actual…
cs.CL2026
From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks
Ayan Datta, Mounika Marreddy, Alexander Mehler +2
Large language models (LLMs) exhibit failures on elementary symbolic tasks such as character counting in a word, despite excelling on complex benchmarks. Although this limitation h…