activity
20242026
collaborators

13 papers

cs.LG2026

ML-Powered LDAP Reconnaissance Detection using Weak Supervision

Shaefer Drew, Edward Raff, Michael Brautbar +6

Lightweight Directory Access Protocol (LDAP) is a protocol that allows users to query and modify Active Directory (AD) data. By default, all users have read access to all AD data t…

cs.CL2026

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang

Large language models are widely adopted as automated evaluation judges, yet the stability of their verdicts under semantically equivalent prompt rephrasings remains largely unexam…

cs.CL2025

Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context

Nilanjana Das, Edward Raff, Aman Chadha +1

As the AI systems become deeply embedded in social media platforms, we've uncovered a concerning security vulnerability that goes beyond traditional adversarial attacks. It becomes…

cs.CL2025

Attribution in Scientific Literature: New Benchmark and Methods

Yash Saxena, Deepa Tilwani, Ali Mohammadi +4

Large language models (LLMs) present a promising yet challenging frontier for automated source citation in scientific communication. Previous approaches to citation generation have…

cs.LG2025

LEACE: Perfect linear concept erasure in closed form

Nora Belrose, David Schneider-Joseph, Shauli Ravfogel +3

Concept erasure aims to remove specified features from an embedding. It can improve fairness (e.g. preventing a classifier from using gender or race) and interpretability (e.g. rem…

cs.CR2025

ClarAVy: A Tool for Scalable and Accurate Malware Family Labeling

Robert J. Joyce, Derek Everett, Maya Fuchs +2

Determining the family to which a malicious file belongs is an essential component of cyberattack investigation, attribution, and remediation. Performing this task manually is time…