10 papers
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
Yiliang Song, Hongjun An, Jiangan Chen +4
Public benchmarks increasingly govern how large language models (LLMs) are ranked, selected, and deployed. We frame this benchmark-centered regime as Silicon Bureaucracy and AI Tes…
Ruyi2 Technical Report
Huan Song, Shuyu Tian, Junyi Hao +5
Large Language Models (LLMs) face significant challenges regarding deployment costs and latency, necessitating adaptive computing strategies. Building upon the AI Flow framework, w…
CreditAudit: 2 Dimension for LLM Evaluation and Selection
Yiliang Song, Hongjun An, Jiangong Xiao +3
Leaderboard scores on public benchmarks have been steadily rising and converging, with many frontier language models now separated by only marginal differences. However, these scor…
Theoretical Foundations of Scaling Law in Familial Models
Huan Song, Qingfei Zhao, Ting Long +4
Neural scaling laws have become foundational for optimizing large language model (LLM) training, yet they typically assume a single dense model output. This limitation effectively…
Single-Pixel Vision-Language Model for Intrinsic Privacy-Preserving Behavioral Intelligence
Hongjun An, Yiliang Song, Jiawei Shao +2
Adverse social interactions, such as bullying, harassment, and other illicit activities, pose significant threats to individual well-being and public safety, leaving profound impac…
Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity
Hongjun An, Yiliang Song, Jiangan Chen +3
Large Language Model (LLM) training often optimizes for preference alignment, rewarding outputs that are perceived as helpful and interaction-friendly. However, this preference-ori…