9 citations · 12 across the 5 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CY2026
Efficient Safety Benchmarking via Item Response Theory
Fabio Spagliardi, Mírian Silva, Ayan Datta +3
Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly…
cs.CL2026
GRAFITE: Generative Regression Analysis Framework for Issue Tracking and Evaluation
Ja Young Lee, Mírian Silva, Mohamed Nasr +6
Large language models (LLMs) are largely motivated by their performance on popular topics and benchmarks at the time of their release. However, over time, contamination occurs due…