9 papers
PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction
Pirzada Suhail, Nagasai Saketh Naidu, Atanu R Sinha +1
Large language models (LLMs) generate text by auto-regressively sampling the next token. This inherently leads to a many-to-many mapping between prompts and responses, complicating…
TIE: A Training-Inversion-Exclusion Framework for Visually Interpretable and Uncertainty-Guided Out-of-Distribution Detection
Pirzada Suhail, Rehna Afroz, Amit Sethi
Deep neural networks often struggle to recognize when an input lies outside their training experience, leading to unreliable and overconfident predictions. Building dependable mach…
EXP-CAM: Explanation Generation and Circuit Discovery Using Classifier Activation Matching
Pirzada Suhail, Aditya Anand, Amit Sethi
Machine learning models, by virtue of training, learn a large repertoire of decision rules for any given input, and any one of these may suffice to justify a prediction. However, i…
Network Inversion for Uncertainty-Aware Out-of-Distribution Detection
Pirzada Suhail, Rehna Afroz, Gouranga Bala +1
Out-of-distribution (OOD) detection and uncertainty estimation (UE) are critical components for building safe machine learning systems, especially in real-world scenarios where une…
Activation Matching for Explanation Generation
Pirzada Suhail, Aditya Anand, Amit Sethi
In this paper we introduce an activation-matching--based approach to generate minimal, faithful explanations for the decision-making of a pretrained classifier on any given image.…
Network Inversion for Generating Confidently Classified Counterfeits
Pirzada Suhail, Pravesh Khaparde, Amit Sethi
In vision classification, generating inputs that elicit confident predictions is key to understanding model behavior and reliability, especially under adversarial or out-of-distrib…