2 papers
cs.CY2026
Efficient Safety Benchmarking via Item Response Theory
Fabio Spagliardi, MÃrian Silva, Ayan Datta +3
Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly…
cs.LG2025
Forecasting MBTA Transit Dynamics: A Performance Benchmarking of Statistical and Machine Learning Models
Sai Siddharth Nalamalpu, Kaining Yuan, Aiden Zhou +1
The Massachusetts Bay Transportation Authority (MBTA) is the main public transit provider in Boston, operating multiple means of transport, including trains, subways, and buses. Ho…