5 papers
Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
Pablo Marcos-Manchón, Rishi Jha, LluÃs Fuentemilla
The Strong Platonic Representation Hypothesis suggests that representational convergence in artificial neural networks can be harnessed constructively: embeddings can be translated…
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
Rishi Jha, Harold Triedman, Arkaprabha Bhattacharya +1
Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. These errors do not thwart ag…
Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
Rishi Jha, Harold Triedman, Justin Wagle +1
Control-flow hijacking attacks manipulate orchestration mechanisms in multi-agent systems into performing unsafe actions that compromise the system and exfiltrate sensitive informa…
Harnessing the Universal Geometry of Embeddings
Rishi Jha, Collin Zhang, Vitaly Shmatikov +1
We introduce the first method for translating text embeddings from one vector space to another without any paired data, encoders, or predefined sets of matches. Our unsupervised ap…
Multi-Agent Systems Execute Arbitrary Malicious Code
Harold Triedman, Rishi Jha, Vitaly Shmatikov
Multi-agent systems coordinate LLM-based agents to perform tasks on users' behalf. In real-world applications, multi-agent systems will inevitably interact with untrusted inputs, s…