activity
20242026
most citedBeyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

Xiao Fei, Yang Zhang, Sarah Almeida Carneiro +1

Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect respo…

cs.CL2026

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

Sarah Almeida Carneiro, Christos Xypolopoulos, Xiao Fei +2

We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-gr…

cs.CL2026

TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models

Yichuan Mo, Yukun Jiang, Yanbo Shi +4

The rapid development of Language Diffusion Models (LDMs) challenges the dominant position of auto-regressive competitors in language processing. However, their flexible, any-order…

cs.CL2026

GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek

Yang Zhang, Mersin Konomi, Christos Xypolopoulos +6

Large Language Models (LLMs) are commonly trained on multilingual corpora that include Greek, yet reliable evaluation benchmarks for Greek-particularly those based on authentic, na…

cs.CL20261 cited

Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning

Yang Zhang, Amr Mohamed, Hadi Abdine +2

Curriculum learning-organizing training data from easy to hard-has improved efficiency across machine learning domains, yet remains underexplored for language model pretraining. We…

cs.CL2025

Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules

Amr Mohamed, Yang Zhang, Michalis Vazirgiannis +1

Diffusion large language models (dLLMs) offer a promising alternative to autoregressive models, but their practical utility is severely hampered by slow, iterative sampling. We pre…