5 citations · 16 across the 16 of their papers we have counts for
14 papers · 1 filter
Narrowing the Gap between Zero- and Few-shot Machine Translation by Matching Styles
Weiting Tan, Haoran Xu, Lingfeng Shen +5
Large language models trained primarily in a monolingual setting have demonstrated their ability to generalize to machine translation using zero- and few-shot examples with in-cont…
Linking Symptom Inventories using Semantic Textual Similarity
Eamonn Kennedy, Shashank Vadlamani, Hannah M Lindsey +79
An extensive library of symptom inventories has been developed over time to measure clinical symptoms, but this variety has led to several long standing issues. Most notably, resul…
MegaWika: Millions of reports and their sources across 50 diverse languages
Samuel Barham, Orion Weller, Michelle Yuan +9
To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages,…
Why Does Zero-Shot Cross-Lingual Generation Fail? An Explanation and a Solution
Tianjian Li, Kenton Murray
Zero-shot cross-lingual transfer is when a multilingual model is trained to perform a task in one language and then is applied to another language. Although the zero-shot cross-lin…
Language Agnostic Code-Mixing Data Augmentation by Predicting Linguistic Patterns
Shuyue Stella Li, Kenton Murray
In this work, we focus on intrasentential code-mixing and propose several different Synthetic Code-Mixing (SCM) data augmentation methods that outperform the baseline on downstream…
Por Qué Não Utiliser Alla Språk? Mixed Training with Gradient Optimization in Few-Shot Cross-Lingual Transfer
Haoran Xu, Kenton Murray
The current state-of-the-art for few-shot cross-lingual transfer learning first trains on abundant labeled data in the source language and then fine-tunes with a few examples on th…