2 papers
cs.AI2026
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
Yiliang Song, Hongjun An, Jiangan Chen +4
Public benchmarks increasingly govern how large language models (LLMs) are ranked, selected, and deployed. We frame this benchmark-centered regime as Silicon Bureaucracy and AI Tes…
cs.CR2025
BlackboxBench: A Comprehensive Benchmark of Black-box Adversarial Attacks
Meixi Zheng, Xuanchen Yan, Zihao Zhu +2
Adversarial examples are well-known tools to evaluate the vulnerability of deep neural networks (DNNs). Although lots of adversarial attack algorithms have been developed, it's sti…