11 papers
Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
Xiao Fei, Yang Zhang, Sarah Almeida Carneiro +1
Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect respo…
CARTE: A Benchmark for Mapping Language Model Knowledge Across France
Sarah Almeida Carneiro, Christos Xypolopoulos, Xiao Fei +2
We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-gr…
Trustworthy Recommendation in the Era of Large Language Models: Opportunities and Challenges
Bohao Wang, Yu Cui, Zhenxiang Xu +13
The field of recommender systems (RS) is currently undergoing two profound paradigm shifts. From the perspective of objectives, the goal has shifted beyond mere recommendation accu…
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
Yichuan Mo, Yukun Jiang, Yanbo Shi +4
The rapid development of Language Diffusion Models (LDMs) challenges the dominant position of auto-regressive competitors in language processing. However, their flexible, any-order…
GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek
Yang Zhang, Mersin Konomi, Christos Xypolopoulos +6
Large Language Models (LLMs) are commonly trained on multilingual corpora that include Greek, yet reliable evaluation benchmarks for Greek-particularly those based on authentic, na…
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
Yang Zhang, Amr Mohamed, Hadi Abdine +2
Curriculum learning-organizing training data from easy to hard-has improved efficiency across machine learning domains, yet remains underexplored for language model pretraining. We…