From the 1 of 5 linked papers with an AI index.
5 papers
ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets
Dipesh Tharu Mahato, Pramod Dhungana
Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not esta…
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal +1
The paper investigates how to suppress specific internal activations in large language models by optimizing only the input prompt, aiming to hide evaluation-awareness signals witho…
Early Warning Signals for OpenVLA Failure under Visual Distribution Shift
Dipesh Tharu Mahato, Rachel Ren
Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activations contain signals associated with th…
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
Dipesh Tharu Mahato
A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards for dual-use biology assistants at the act…
TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift
Dipesh Tharu Mahato, Rohan Poudel, Pramod Dhungana
Deep neural networks often achieve high accuracy, but ensuring their reliability under adversarial and distributional shifts remains a pressing challenge. We propose TriGuard, a un…