156 citations · 156 across the 8 of their papers we have counts for
1 paper · 1 filter
Masaharu Mizumoto, Dat Nguyen, Zhiheng Han +3
Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarit…