contextual testing 1evaluation workflow 1human-aligned evaluation 1reliability gating 1rubric-based scoring 1
From the 1 of 5 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
Gabriel Chua, Leanne Tan, Ziyu Ge +1
Large language models (LLMs) often fail to maintain safety in low-resource language varieties, such as code-mixed vernaculars and regional dialects. We introduce RabakBench, a mult…
cs.CL2025
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
Leanne Tan, Gabriel Chua, Ziyu Ge +1
Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments…
cs.CL2025
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
Ziyu Ge, Gabriel Chua, Leanne Tan +1
As online communication increasingly incorporates under-represented languages and colloquial dialects, standard translation systems often fail to preserve local slang, code-mixing,…