collaborators

12 papers

cs.CR2026

SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents

Rui Melo, Riccardo Fogliato, Sean Zhou +2

Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is merged into shared repositories. However,…

cs.AI2026

Personalization, Personas, and Forecasting in Value Alignment

James Wedgwood, Pratiksha Thaker, Neil Kale +1

LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden quest…

cs.CY2026

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety

Neil Kale, Rebecca Portnoff, Pratiksha Thaker +5

Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child sexual abuse material, facilit…

cs.LG2025

Membership Inference Attacks for Unseen Classes

Pratiksha Thaker, Neil Kale, Zhiwei Steven Wu +1

A key tool in developing safe AI models is \emph{data auditing}, i.e., using statistical tools to determine whether harmful content may have been used in the training data of a bla…

cs.LG2025

PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries

Steven Kolawole, Keshav Santhanam, Virginia Smith +1

LLM serving systems typically treat user prompts as monolithic inputs, optimizing inference through decoding tricks or inter-query batching. However, many real-world prompts contai…

cs.LG2025

On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift

Pratiksha Thaker, Amrith Setlur, Zhiwei Steven Wu +1

Public pretraining is a promising approach to improve differentially private model training. However, recent work has noted that many positive research results studying this paradi…