2 papers
stat.ML2026
An Interpretable and Scalable Framework for Evaluating Large Language Models
Xinhao Qu, Qiang Heng, Hao Zeng +1
Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the inherent stochasticity of LLM…
stat.ME2026
Transfer Learning for Robust Structured Regression with Bi-level Source Detection
Haoming Shi, Yang Feng, Xiaoqian Liu
High-dimensional data in modern applications, such as COVID-19 mortality, often span multiple domains. Leveraging auxiliary information from source domains to improve performance i…