6 papers
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
Punya Syon Pandey, Hai Son Le, Devansh Bhardwaj +2
Large language models (LLMs) are increasingly deployed in contexts where their failures can have direct sociopolitical consequences. Yet, existing safety benchmarks rarely test vul…
Adaptive Estimation of Drifting Noise in Quantum Error Correction
Devansh Bhardwaj, Evangelia Takou, Yingjia Lin +1
Advancing quantum information processors and building fault-tolerant architectures rely on the ability to accurately characterize the noise sources and suppress their impact on qua…
Agent Context Protocols Enhance Collective Inference
Devansh Bhardwaj, Arjun Beniwal, Shreyas Chaudhari +5
AI agents have become increasingly adept at complex tasks such as coding, reasoning, and multimodal understanding. However, building generalist systems requires moving beyond indiv…
Accelerated Smoothing: A Scalable Approach to Randomized Smoothing
Devansh Bhardwaj, Kshitiz Kaushik, Sarthak Gupta
Randomized smoothing has emerged as a potent certifiable defense against adversarial attacks by employing smoothing noises from specific distributions to ensure the robustness of a…
Invisible Traces: Using Hybrid Fingerprinting to identify underlying LLMs in GenAI Apps
Devansh Bhardwaj, Naman Mishra
Fingerprinting refers to the process of identifying underlying Machine Learning (ML) models of AI Systemts, such as Large Language Models (LLMs), by analyzing their unique characte…
Rethinking Randomized Smoothing from the Perspective of Scalability
Anupriya Kumari, Devansh Bhardwaj, Sukrit Jindal
Machine learning models have demonstrated remarkable success across diverse domains but remain vulnerable to adversarial attacks. Empirical defense mechanisms often fail, as new at…