4 papers
High-Dimensional Importance-Weighted Information Criteria: Theory and Optimality
Yong-Syun Cao, Shinpei Imori, Ching-Kang Ing
Imori and Ing (2025) proposed the importance-weighted orthogonal greedy algorithm (IWOGA) for model selection in high-dimensional misspecified regression models under covariate shi…
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
Fengze Liu, Weidong Zhou, Binbin Liu +8
Quality and diversity are two critical metrics for the training data of large language models (LLMs), positively impacting performance. Existing studies often optimize these metric…
Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge
Li Zhou, Taelin Karidi, Wanlong Liu +5
Recent studies have highlighted the presence of cultural biases in Large Language Models (LLMs), yet often lack a robust methodology to dissect these phenomena comprehensively. Our…
Rethinking the Effectiveness of Graph Classification Datasets in Benchmarks for Assessing GNNs
Zhengdao Li, Yong Cao, Kefan Shuai +2
Graph classification benchmarks, vital for assessing and developing graph neural networks (GNNs), have recently been scrutinized, as simple methods like MLPs have demonstrated comp…