4 papers
ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets
Dipesh Tharu Mahato, Pramod Dhungana
Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not esta…
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
Dipesh Tharu Mahato
A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards for dual-use biology assistants at the act…
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal +1
Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal la…
TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift
Dipesh Tharu Mahato, Rohan Poudel, Pramod Dhungana
Deep neural networks often achieve high accuracy, but ensuring their reliability under adversarial and distributional shifts remains a pressing challenge. We propose TriGuard, a un…