3 citations · 3 across the 4 of their papers we have counts for
4 papers · 1 filter
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
Yunsu Kim, Kaden Uhlig, Ashwin Purohit +10
Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 a…
GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation
Yunsu Kim, Kaden Uhlig, Joern Wuebker
Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editing. We argue that, for agenti…
Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization
Kaden Uhlig, Joern Wuebker, Raphael Reinauer +1
Reinforcement Learning from Human Feedback (RLHF) and derivative techniques like Direct Preference Optimization (DPO) are task-alignment algorithms used to repurpose general, found…
Neural Machine Translation Models Can Learn to be Few-shot Learners
Raphael Reinauer, Patrick Simianer, Kaden Uhlig +2
The emergent ability of Large Language Models to use a small number of examples to learn to perform in novel domains and tasks, also called in-context learning (ICL). In this work,…