16 papers
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Mubashara Akhtar, Anka Reuel, Prajna Soni +36
Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult…
AdaMame: A Training Recipe for Adaptive Multilingual Reasoning
Dayeon Ki, Kevin Duh, Marine Carpuat
While Large Reasoning Models (LRMs) show strong performance in English, they often fail to reason in the language of the query, a phenomenon known as language collapse. Existing RL…
Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
Dayeon Ki, Marine Carpuat, Paul McNamee +4
Multilingual Retrieval-Augmented Generation (mRAG) systems enable language models to answer knowledge-intensive queries with citation-supported responses across languages. Despite…
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
Dayeon Ki, Yu Hou, Rachel Rudinger +3
Neologisms and emerging slang are central to daily conversation, yet challenging for non-native speakers (NNS) to interpret and use appropriately in cross-cultural communication wi…
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
Dayeon Ki, Kevin Duh, Marine Carpuat
Large Reasoning Models (LRMs) still exhibit large performance gaps between English and other languages, yet much current work assumes these gaps can be closed simply by making reas…
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
Lingjun Zhao, Dayeon Ki, Marine Carpuat +1
Language models are known to exhibit various forms of cultural bias in decision-making tasks, yet much less is known about their degree of cultural familiarity in open-ended text g…