5 papers
Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment -- Is English Enough?
Adnan Al Ali, Kathy Hämmerl, Kathy Hämmerl +3
Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to E…
Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking
Songbo Hu, Yinhong Liu, Ej Zhou +5
Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-parallel datasets at scale. This…
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
Felix Friedrich, Katharina Hämmerl, Patrick Schramowski +4
Text-to-image generation models have recently achieved astonishing results in image quality, flexibility, and text alignment, and are consequently employed in a fast-growing number…
Beyond Literal Token Overlap: Token Alignability for Multilinguality
Katharina Hämmerl, Tomasz Limisiewicz, JindÅich Libovický +1
Previous work has considered token overlap, or even similarity of token distributions, as predictors for multilinguality and cross-lingual knowledge transfer in language models. Ho…
Understanding Cross-Lingual Alignment -- A Survey
Katharina Hämmerl, JindÅich Libovický, Alexander Fraser
Cross-lingual alignment, the meaningful similarity of representations across languages in multilingual language models, has been an active field of research in recent years. We sur…