4 citations · 7 across the 3 of their papers we have counts for
5 papers · 1 filter
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
Yunsu Kim, Kaden Uhlig, Ashwin Purohit +10
Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 a…
Neural Machine Translation Models Can Learn to be Few-shot Learners
Raphael Reinauer, Patrick Simianer, Kaden Uhlig +2
The emergent ability of Large Language Models to use a small number of examples to learn to perform in novel domains and tasks, also called in-context learning (ICL). In this work,…
STAR: A Schema-Guided Dialog Dataset for Transfer Learning
Johannes E. M. Mosig, Shikib Mehri, Thomas Kober
We present STAR, a schema-guided task-oriented dialog dataset consisting of 127,833 utterances and knowledge base queries across 5,820 task-oriented dialogs in 13 domains that is e…
Where is the context? -- A critique of recent dialogue datasets
Johannes E. M. Mosig, Vladimir Vlasov, Alan Nichol
Recent dialogue datasets like MultiWOZ 2.1 and Taskmaster-1 constitute some of the most challenging tasks for present-day dialogue models and, therefore, are widely used for system…
Dialogue Transformers
Vladimir Vlasov, Johannes E. M. Mosig, Alan Nichol
We introduce a dialogue policy based on a transformer architecture, where the self-attention mechanism operates over the sequence of dialogue turns. Recent work has used hierarchic…