5 papers
Rethinking Cross-lingual Gaps from a Statistical Viewpoint
Vihari Piratla, Purvam Jain, Darshan Singh +3
Any piece of knowledge is usually expressed in one or a handful of natural languages on the web or in any large corpus. Large Language Models (LLMs) act as a bridge by acquiring kn…
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji +2
Culturally aware safeguards are crucial for AI alignment in real-world settings, where safety extends beyond common sense and encompasses diverse local values, norms, and region-sp…
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji +2
Safeguard models help large language models (LLMs) detect and block harmful content, but most evaluations remain English-centric and overlook linguistic and cultural diversity. Exi…
LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
Alham Fikri Aji, Trevor Cohn
As one of the world's most populous countries, with 700 languages spoken, Indonesia is behind in terms of NLP progress. We introduce LoraxBench, a benchmark that focuses on low-res…
Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation
Jianyuan Guo, Peike Li, Trevor Cohn
Sign Language Translation (SLT) aims to map sign language videos to spoken language text. A common approach relies on gloss annotations as an intermediate representation, decomposi…