1 citations · 1 across the 6 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
Marek Jeliński, Jan Dubiński, Maciej Chrabaszcz +1
Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated t…
cs.AI2026
The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection
Wojciech Zarzecki, Jan Dubiński, Sebastian Cygert
Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for detecting training-data member…