17 citations · 35 across the 11 of their papers we have counts for
13 papers · 1 filter
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
Yunsu Kim, Kaden Uhlig, Ashwin Purohit +10
Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 a…
GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation
Yunsu Kim, Kaden Uhlig, Joern Wuebker
Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editing. We argue that, for agenti…
Self-Training with Purpose Preserving Augmentation Improves Few-shot Generative Dialogue State Tracking
Jihyun Lee, Chaebin Lee, Yunsu Kim +1
In dialogue state tracking (DST), labeling the dataset involves considerable human labor. We propose a new self-training framework for few-shot generative DST that utilize unlabele…
Multi-Type Conversational Question-Answer Generation with Closed-ended and Unanswerable Questions
Seonjeong Hwang, Yunsu Kim, Gary Geunbae Lee
Conversational question answering (CQA) facilitates an incremental and interactive understanding of a given context, but building a CQA system is difficult for many domains due to…
When and Why is Unsupervised Neural Machine Translation Useless?
Yunsu Kim, Miguel Graça, Hermann Ney
This paper studies the practicality of the current state-of-the-art unsupervised methods in neural machine translation (NMT). In ten translation tasks with various data settings, w…
When and Why is Document-level Context Useful in Neural Machine Translation?
Yunsu Kim, Duc Thanh Tran, Hermann Ney
Document-level context has received lots of attention for compensating neural machine translation (NMT) of isolated sentences. However, recent advances in document-level NMT focus…