collaborators

31 papers

cs.CL2026

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

Xiao Fei, Yang Zhang, Sarah Almeida Carneiro +1

Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect respo…

cs.HC2026

From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Agent

Mingyu Huang, Weiqing Min, Ying Jin +2

Personalized glucose regulation remains a central yet unresolved challenge in precision nutrition, as postprandial glucose response varies substantially across individuals. Existin…

cs.CL2026

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

Yang Zhang, Xiao Fei, Amr Mohamed +6

Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural knowledge is better accessed thr…

cs.CL2026

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

Sarah Almeida Carneiro, Christos Xypolopoulos, Xiao Fei +2

We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-gr…

cs.LG2026

A Kinetic Energy Perspective of Flow Matching

Ziyun Li, Huancheng Hu, Soon Hoe Lim +6

Flow-based generative models can be viewed through a physics lens: sampling transports a particle from noise to data by integrating a learned velocity field, and each sample corres…

cs.LG2026

GraViti: Graph-Level Variational Autoencoders with Relaxed Permutation Invariance

Roman Bresson, Konstantinos Divriotis, Johannes F. Lutzeyer +2

We introduce GraViti, a transformer-based graph-level variational autoencoder that maps entire graphs to compact latent vectors. This design produces a true graph-level latent spac…