5 papers
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
Eeham Khan, Firas Saidani, Owen Van Esbroeck +2
Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain largely confined to a small number of high-resource languages for which there i…
Generating High-quality Privacy-preserving Synthetic Data
David Yavo, Richard Khoury, Christophe Pere +1
Synthetic tabular data enables sharing and analysis of sensitive records, but its practical deployment requires balancing distributional fidelity, downstream utility, and privacy p…
Neural Machine Translation for Coptic-French: Strategies for Low-Resource Ancient Languages
Nasma Chaoui, Richard Khoury
This paper presents the first systematic study of strategies for translating Coptic into French. Our comprehensive pipeline systematically evaluates: pivot versus direct translatio…
Automated Journalistic Questions: A New Method for Extracting 5W1H in French
Maxence Verhaverbeke, Julie A. Gramaccia, Richard Khoury
The 5W1H questions -- who, what, when, where, why and how -- are commonly used in journalism to ensure that an article describes events clearly and systematically. Answering them i…
Preference-based learning for news headline recommendation
Alexandre Bouras, Audrey Durand, Richard Khoury
This study explores strategies for optimizing news headline recommendations through preference-based learning. Using real-world data of user interactions with French-language onlin…