4 papers · 1 filter
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
Nikolay Banar, Ehsan Lotfi, Jens Van Nooten +3
Recently, embedding resources, including models, benchmarks, and datasets, have been widely released to support a variety of languages. However, the Dutch language remains underrep…
BEIR-NL: Zero-shot Information Retrieval Benchmark for the Dutch Language
Nikolay Banar, Ehsan Lotfi, Walter Daelemans
Zero-shot evaluation of information retrieval (IR) models is often performed using BEIR; a large and heterogeneous benchmark composed of multiple datasets, covering different retri…
Bilingual BSARD: Extending Statutory Article Retrieval to Dutch
Ehsan Lotfi, Nikolay Banar, Nerses Yuzbashyan +1
Statutory article retrieval plays a crucial role in making legal information more accessible to both laypeople and legal professionals. Multilingual countries like Belgium present…
Character-level Transformer-based Neural Machine Translation
Nikolay Banar, Walter Daelemans, Mike Kestemont
Neural machine translation (NMT) is nowadays commonly applied at the subword level, using byte-pair encoding. A promising alternative approach focuses on character-level translatio…