2 papers
cs.LG2026
Transformers as Unsupervised Learning Algorithms: A study on Gaussian Mixtures
Zhiheng Chen, Ruofan Wu, Guanhua Fang
The transformer architecture has demonstrated remarkable capabilities in modern artificial intelligence, among which the capability of implicitly learning an internal model during…
cs.AI2025
Toward a unified framework for data-efficient evaluation of large language models
Lele Liao, Qile Zhang, Ruofan Wu +1
Evaluating large language models (LLMs) on comprehensive benchmarks is a cornerstone of their development, yet it's often computationally and financially prohibitive. While Item Re…