29 citations · 255 across the 117 of their papers we have counts for
128 papers
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
EDRAC: Benchmarking Arabic Dialect Reading Comprehension
Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn +15
Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing A…
DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer
Patomporn Payoungkhamdee, Tinnakit Udsa, Jian Gang Ngui +3
Small language models (SLMs) are efficient and scalable, but their multilingual capabilities degrade severely at sub-billion scales, especially for Southeast Asian (SEA) languages.…
TimpaTeks: Automatic In-place Text Sequence Modification via Diffusion Language Model Steering
Ryandito Diandaru, Ikhlasul Akmal Hanif, Fadli Aulawi Al Ghiffari +2
We extend activation steering to diffusion language models (DLMs) and study a novel problem that arose due to the inference mechanism of DLMs: Modifying a text in-place to manifest…
SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding
Peerawat Chomphooyod, Jian Gang Ngui, Yosephine Susanto +5
Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NLI benchmarks are largely Wes…
Sense Representations Are Inducible Interfaces
Jan Christian Blaise Cruz, Alham Fikri Aji
Sense representations (explicit, per-token meaning decompositions) are useful for disambiguation, steering, and cross-lingual alignment, but existing approaches require models to b…