5 papers
LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents
Omar El Bachyr, Fred Philippy, Laura Maria Bernardy +3
Recent page-image retrievers such as ColPali have improved retrieval over visually rich documents, yet little is known about how they behave in cross-lingual, low-resource settings…
ltzGLUE: Luxembourgish General Language Understanding Evaluation
Alistair Plum, Felicia Körner, Anne-Marie Lutgen +8
This paper presents ltzGLUE, the first Natural Language Understanding (NLU) benchmark for Luxembourgish (LTZ) based on the popular GLUE benchmark for English. Although NLU tasks ar…
Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
Alistair Plum, Laura Bernardy, Tharindu Ranasinghe
We present judgeWEL, a dataset for named entity recognition (NER) in Luxembourgish, automatically labelled and subsequently verified using large language models (LLM) in a novel pi…
LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
Fred Philippy, Laura Bernardy, Siwen Guo +2
Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human prompts. However, low-resource languages…
POV Learning: Individual Alignment of Multimodal Models using Human Perception
Simon Werner, Katharina Christ, Laura Bernardy +2
Aligning machine learning systems with human expectations is mostly attempted by training with manually vetted human behavioral samples, typically explicit feedback. This is done o…