From the 1 of 27 linked papers with an AI index.
27 papers
Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Osvaldo Quinjica, Eric Bennett, Xinchen Yang +2
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evalua…
Contrastive ESA: Human Evaluation of Multiple Translations at Once
Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee +6
The paper proposes Contrastive Error Span Annotation (cESA), a human evaluation protocol that shows multiple translations of the same source together, lets annotators mark error sp…
Measuring User's Mental Models of Speech Translation in Human-AI Collaboration
HyoJung Han, Nishant Balepur, Jordan Boyd-Graber +1
Millions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do. This paper studies users' mental models o…
AdaMame: A Training Recipe for Adaptive Multilingual Reasoning
Dayeon Ki, Kevin Duh, Marine Carpuat
While Large Reasoning Models (LRMs) show strong performance in English, they often fail to reason in the language of the query, a phenomenon known as language collapse. Existing RL…
Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
Dayeon Ki, Marine Carpuat, Paul McNamee +4
Multilingual Retrieval-Augmented Generation (mRAG) systems enable language models to answer knowledge-intensive queries with citation-supported responses across languages. Despite…
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios
Giuseppe Attanasio, Beatrice Savoldi, Daniel Chechelnitsky +4
Speech translation (ST) is increasingly adopted in user applications, yet its evaluation largely focuses on decontextualized testbeds and holistic quality, rather than end users' c…