4 papers
PRISM: Recovering Instruction Sets from Language Model Activations
Gilad Gressel, Rahul Pankajakshan, Julia Diament +3
As LLMs are deployed as agents, reliable monitoring requires knowing not only what they output, but which instructions are steering their behavior. This is difficult when models in…
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
Shir Rozenfeld, Rahul Pankajakshan, Itay Zloczower +3
Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. Ho…
Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
Gilad Gressel, Rahul Pankajakshan, Shir Rozenfeld +4
Romance-baiting scams have become a major source of financial and emotional harm worldwide. These operations are run by organized crime syndicates that traffic thousands of people…
Are You Human? An Adversarial Benchmark to Expose LLMs
Gilad Gressel, Rahul Pankajakshan, Yisroel Mirsky
Large Language Models (LLMs) have demonstrated an alarming ability to impersonate humans in conversation, raising concerns about their potential misuse in scams and deception. Huma…