works on

From the 1 of 8 linked papers with an AI index.

collaborators

8 papers

cs.CL2026

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

Ej Zhou, Lucas Resck, Zheng Hui +1

The paper shows that large language model evaluators give systematically different scores to the same content in different languages, favoring lower‑resource languages, even though…

cs.CL2026

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…

cs.CL2026

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

Songbo Hu, Yinhong Liu, Ej Zhou +5

Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-parallel datasets at scale. This…

cs.LG2026

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

Jiaqi Weng, Han Zheng, Hanyu Zhang +6

Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, under what circumstances SAEs derive mos…

cs.CY2026

Artificial intelligence is creating a new global linguistic hierarchy

Giulia Occhini, Kumiko Tanaka-Ishii, Anna Barford +9

Artificial intelligence (AI) has the potential to transform healthcare, education, governance and socioeconomic equity, but its benefits remain concentrated in a small number of la…

cs.CL2025

Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models

Ej Zhou, Caiqi Zhang, Tiancheng Hu +4

Confidence calibration, the alignment of a model's predicted confidence with its actual accuracy, is crucial for the reliable deployment of Large Language Models (LLMs). However, t…