11 papers
Unsupervised Features Mining via Activation Geometry
Amit LeVi, Elad David, Max Fomin
Interpretability methods aim to reveal the features represented inside large language models (LLMs). Many existing methods begin with labeled examples of a human-defined concept th…
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
Dvir Alsheich, Adar Peleg, Ben Hagag +3
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents collaborate to execute complex workflows. Effe…
Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring
Max Fomin, Elad David, Amit LeVi
Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated. We ask when an internal readout suppo…
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
Amit LeVi, Raz Lapid, Rom Himelstein +3
Many LLM applications require only narrow capabilities, yet standard post-training quantization (PTQ) methods allocate precision without considering the target task. This can waste…
Mirage Probes: How Vision Models Fake Visual Understanding
Daniel Ben-Levi, Judah Goldfeder, Weiliang Zhao +5
Vision-language models (VLMs) can answer image-based questions confidently, and often correctly, even when no image is provided. This mirage behavior inflates benchmark scores with…
Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software
Tomer Kordonsky, Amit LeVi, Maayan Yamin +2
LLMs are increasingly used for code generation, but their outputs often follow recurring templates that can induce predictable vulnerabilities. We study vulnerability persistence i…