Publications (9)
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
Yaniv Nikankin, Dana Arad, Itay Itzhak +4
One of the main challenges in mechanistic interpretability is circuit discovery, determining which parts of a model perform a given task. We build on the Mechanistic Interpretabili…
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
Adi Simhi, Jonathan Herzig, Martin Tutek +3
As large language models (LLMs) evolve from conversational assistants into autonomous agents, evaluating the safety of their actions becomes critical. Prior safety benchmarks have…
HACK: Hallucinations Along Certainty and Knowledge Axes
Adi Simhi, Jonathan Herzig, Itay Itzhak +7
Hallucinations in LLMs present a critical barrier to their reliable usage. Existing research usually categorizes hallucination by their external properties rather than by the LLMs'…
Interpreting Embedding Spaces by Conceptualization
Adi Simhi, Shaul Markovitch
One of the main methods for computational interpretation of a text is mapping it into a vector in some embedding space. Such vectors can then be used for a variety of textual proce…
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, :, Teven Le Scao +391
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
Adi Simhi, Itay Itzhak, Fazl Barez +2
Prior work on large language model (LLM) hallucinations has associated them with model uncertainty or inaccurate knowledge. In this work, we define and investigate a distinct type…
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
Adi Simhi, Fazl Barez, Martin Tutek +2
How does the conversational past of large language models (LLMs) influence their future performance? Recent work suggests that LLMs are affected by their conversational history in…
Distinguishing Ignorance from Error in LLM Hallucinations
Adi Simhi, Jonathan Herzig, Idan Szpektor +1
Large language models (LLMs) are susceptible to hallucinations -- factually incorrect outputs -- leading to a large body of work on detecting and mitigating such cases. We argue th…
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
Adi Simhi, Jonathan Herzig, Idan Szpektor +1
Large language models (LLMs) are prone to hallucinations, which sparked a widespread effort to detect and prevent them. Recent work attempts to mitigate hallucinations by interveni…