3 papers
cs.AI2026
Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
Vassilis Papadopoulos, McNair Shah, Sam Zimmerman +1
AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of…
cs.AI2026
Me, Myself, and : Evaluating and Explaining LLM Introspection
Atharv Naphade, Samarth Bhargav, Sean Lim +1
A hallmark of human intelligence is Introspection-the ability to assess and reason about one's own cognitive processes. Introspection has emerged as a promising but contested capab…
cs.AI2025
The Geometry of Harmfulness in LLMs through Subconcept Probing
McNair Shah, Saleena Angeline, Adhitya Rajendra Kumar +5
Recent advances in large language models (LLMs) have intensified the need to understand and reliably curb their harmful behaviours. We introduce a multidimensional framework for pr…