1 citations · 1 across the 17 of their papers we have counts for
18 papers · 1 filter
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca +5
Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in Eng…
Dynamically Allocating Evaluation Effort for Model Ranking
Vilém Zouhar, Julia Kreutzer, Alon Lavie +4
While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation pr…
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
Eduardo Sánchez, Rita Berrada, Dan-Mircea Mirea +8
Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the sys…
JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
Erlis Lushtaku, Bora Kargi, Ali Elganzory +4
LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a spe…
The Culture Funnel: You Can't Align What isn't in the Data
Ananya Sahu, Mehrnaz Mofakhami, Daniel D'Souza +3
Current cultural alignment approaches focus on inference-time interventions, assuming models already contain sufficient cultural knowledge. We argue modern LLM pipelines suffer fro…
Tiny Aya: Bridging Scale and Multilingual Depth
Alejandro R. Salamanca, Diana Abagyan, Daniel D'souza +23
Tiny Aya redefines what a small multilingual language model can achieve. Trained on 70 languages and refined through region-aware posttraining, it delivers state-of-the-art in tran…