collaborators

5 papers

cs.CL2026

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

Andrea Caciolai, Pere-Lluís Huguet Cabot, Chierh Cheng +11

Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI…

cs.CL2026

The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

Piotr Nawrot, Robert Li, Renjie Huang +3

Sparse attention offers a promising strategy to extend long-context capabilities in Transformer LLMs, yet its efficiency-accuracy trade-offs remain unclear due to the lack of compr…

cs.CL2026

Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages

Anaelia Ovalle, Candace Ross, Sebastian Ruder +4

Large language models demonstrate strong reasoning capabilities through chain-of-thought prompting, but whether this reasoning quality transfers across languages remains underexplo…

cs.CL2025

Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks

Bhaktipriya Radharapu, Manon Revel, Megan Ung +2

The increasing use of LLMs as substitutes for humans in ``aligning'' LLMs has raised questions about their ability to replicate human judgments and preferences, especially in ambiv…

cs.CL2025

AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic

Nathaniel R. Robinson, Shahd Abdelmoneim, Kelly Marchisio +1

Dialectal Arabic (DA) varieties are under-served by language technologies, particularly large language models (LLMs). This trend threatens to exacerbate existing social inequalitie…