4 citations · 8 across the 11 of their papers we have counts for
Showing 2026 · cs.AIShow all
2 papers · 2 filters
cs.AI2026
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
Yiliang Song, Hongjun An, Jiangan Chen +4
Public benchmarks increasingly govern how large language models (LLMs) are ranked, selected, and deployed. We frame this benchmark-centered regime as Silicon Bureaucracy and AI Tes…
cs.AI2026
CreditAudit: 2 Dimension for LLM Evaluation and Selection
Yiliang Song, Hongjun An, Jiangong Xiao +3
Leaderboard scores on public benchmarks have been steadily rising and converging, with many frontier language models now separated by only marginal differences. However, these scor…