collaborators

9 papers

cs.AI2026

Code Monitor Red Teaming for Public-Test-Passing Code

Junchi Liao, Jiawen Deng, Fuji Ren

Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has p…

cs.AI2026

Personalization, Personas, and Forecasting in Value Alignment

James Wedgwood, Pratiksha Thaker, Neil Kale +1

LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden quest…

cs.CY2026

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety

Neil Kale, Rebecca Portnoff, Pratiksha Thaker +5

Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child sexual abuse material, facilit…

cs.CR2026

Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders

David Campbell, Neil Kale, Udari Madhushani Sehwag +5

Safety alignment in large language models (LLMs), particularly for cybersecurity tasks, primarily focuses on preventing misuse. While this approach reduces direct harm, it obscures…

cs.LG2025

Membership Inference Attacks for Unseen Classes

Pratiksha Thaker, Neil Kale, Zhiwei Steven Wu +1

A key tool in developing safe AI models is \emph{data auditing}, i.e., using statistical tools to determine whether harmful content may have been used in the training data of a bla…

cs.AI2025

Mitigating Modal Imbalance in Multimodal Reasoning

Chen Henry Wu, Neil Kale, Aditi Raghunathan

Foundation models (FMs) deployed in real-world tasks such as computer-use agents must integrate diverse modalities. How good are FMs at performing joint reasoning, simultaneously r…