4 papers
You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring
Aritra Das, Jaee Ponde, Mihir More +1
LLM agents are increasingly deployed in multi-agent systems, where they can collude while keeping their actions benign. Output monitors designed to detect such collusions can be fo…
Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures
Jaee Ponde, Roshni Agarwal, Subhashis Banerjee
Neural networks are increasingly employed to identify both well-defined and ambiguous concepts, yet output-level metrics reveal little about how those concepts are represented inte…
On the Indistinguishability of Human v/s AI Generated Text
Jaee Ponde, Aritra Das, Mihir More +1
The rapid improvement of LLMs has made distinguishing AI-generated text from human writing a pressing problem. This challenge is further amplified by paraphrasing tools designed to…
Does Order Matter : Connecting The Law of Robustness to Robust Generalization
Mihir More, Aritra Das, Jaee Ponde +3
Bubeck and Selke (2021) propose the connection between the Law of Robustness and robust generalization error as an open problem. The Law of Robustness states that overparameterizat…