3 papers
cs.LG2026
From token probabilities to calibrated confidence: An empirical study of mathematical question answering
Avery Ma, Lorne Schell, Vin Bhaskara +1
Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns these estimates with empirical…
cs.CR2026
StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection
Zhuoxin Zhan, Akbar Rafiey, Avery Ma +2
Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we…
cs.CL2025
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
Avery Ma, Yangchen Pan, Amir-massoud Farahmand
Many-shot jailbreaking circumvents the safety alignment of LLMs by exploiting their ability to process long input sequences. To achieve this, the malicious target prompt is prefixe…