Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Unsupervised Features Mining via Activation Geometry
Amit LeVi, Elad David, Max Fomin
Interpretability methods aim to reveal the features represented inside large language models (LLMs). Many existing methods begin with labeled examples of a human-defined concept th…
cs.AI2026
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
Dvir Alsheich, Adar Peleg, Ben Hagag +3
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents collaborate to execute complex workflows. Effe…