works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.AI2026

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

Leanne Tan, Rohan Jaggi, Shaun Khoo +1

The paper introduces Kaleidoscope, an integrated workflow that combines persona‑based test generation, contextual rubrics, and human review with LLM‑based judges to provide reliabl…

cs.CL2026

Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps

Gabriel Chua, Leanne Tan, Ziyu Ge +1

Large language models (LLMs) often fail to maintain safety in low-resource language varieties, such as code-mixed vernaculars and regional dialects. We introduce RabakBench, a mult…

cs.CL2025

LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators

Leanne Tan, Gabriel Chua, Ziyu Ge +1

Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments…

cs.CL2025

Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation

Ziyu Ge, Gabriel Chua, Leanne Tan +1

As online communication increasingly incorporates under-represented languages and colloquial dialects, standard translation systems often fail to preserve local slang, code-mixing,…

cs.SE2025

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications

Jia Yi Goh, Shaun Khoo, Nyx Iskandar +3

Most safety testing efforts for large language models (LLMs) today focus on evaluating foundation models. However, there is a growing need to evaluate safety at the application lev…