4 papers · 1 filter
Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
Xiao Fei, Yang Zhang, Sarah Almeida Carneiro +1
Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect respo…
The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs
Yang Zhang, Xiao Fei, Amr Mohamed +6
Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural knowledge is better accessed thr…
CARTE: A Benchmark for Mapping Language Model Knowledge Across France
Sarah Almeida Carneiro, Christos Xypolopoulos, Xiao Fei +2
We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-gr…
Graph Linearization Methods for Reasoning on Graphs with Large Language Models
Christos Xypolopoulos, Guokan Shang, Xiao Fei +6
Large language models have evolved to process multiple modalities beyond text, such as images and audio, which motivates us to explore how to effectively leverage them for graph re…