4 papers
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
Hour Kaing, Raj Dabre, Haiyue Song +3
This work introduces {\it PrahokBART}, a compact pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora. We focus o…
Structured Document Translation via Format Reinforcement Learning
Haiyue Song, Johannes Eschbach-Dymanus, Hour Kaing +4
Recent works on structured text translation remain limited to the sentence level, as they struggle to effectively handle the complex document-level XML or HTML structures. To addre…
When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
Ahmed Elshabrawy, Hour Kaing, Haiyue Song +4
Alignment with high-resource standard languages is often assumed to aid the modeling of related low-resource varieties. We challenge this assumption by demonstrating that excessive…
Connecting Ideas in 'Lower-Resource' Scenarios: NLP for National Varieties, Creoles and Other Low-resource Scenarios
Aditya Joshi, Diptesh Kanojia, Heather Lent +2
Despite excellent results on benchmarks over a small subset of languages, large language models struggle to process text from languages situated in `lower-resource' scenarios such…