collaborators

5 papers

cs.CL2026

"Be My Cheese?": Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs

Madison Van Doren, Casey Ford, Jennifer Barajas +2

We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilingual large language models (LLMs…

cs.CL2026

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

Casey Ford, Madison Van Doren, Sicheng Jin +1

The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures. We present the first syste…

cs.CL2026

Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases

Casey Ford, Madison Van Doren, Emily Dix

Multimodal large language models (MLLMs) are increasingly deployed in real-world systems, yet their safety under adversarial prompting remains underexplored. We present a two-phase…

cs.CL2025

Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models

Madison Van Doren, Casey Ford

Multimodal large language models (MLLMs) are increasingly used in real world applications, yet their safety under adversarial conditions remains underexplored. This study evaluates…

cs.CL2025

"Be My Cheese?": Assessing Cultural Nuance in Multilingual LLM Translations

Madison Van Doren, Cory Holland

This pilot study explores the localisation capabilities of state-of-the-art multilingual AI models when translating figurative language, such as idioms and puns, from English into…