papers

Publications (9)

cs.CL2025

BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection

Yaniv Nikankin, Dana Arad, Itay Itzhak +4

One of the main challenges in mechanistic interpretability is circuit discovery, determining which parts of a model perform a given task. We build on the Mechanistic Interpretabili…

cs.CL2026

ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs

Adi Simhi, Jonathan Herzig, Martin Tutek +3

As large language models (LLMs) evolve from conversational assistants into autonomous agents, evaluating the safety of their actions becomes critical. Prior safety benchmarks have…

cs.CL2025

HACK: Hallucinations Along Certainty and Knowledge Axes

Adi Simhi, Jonathan Herzig, Itay Itzhak +7

Hallucinations in LLMs present a critical barrier to their reliable usage. Existing research usually categorizes hallucination by their external properties rather than by the LLMs'…

cs.CL2023

Interpreting Embedding Spaces by Conceptualization

Adi Simhi, Shaul Markovitch

One of the main methods for computational interpretation of a text is mapping it into a vector in some embedding space. Such vectors can then be used for a variety of textual proce…

cs.CL2023

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao +391

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…

cs.CL2025

Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Adi Simhi, Itay Itzhak, Fazl Barez +2

Prior work on large language model (LLM) hallucinations has associated them with model uncertainty or inaccurate knowledge. In this work, we define and investigate a distinct type…

cs.CL2026

Old Habits Die Hard: How Conversational History Geometrically Traps LLMs

Adi Simhi, Fazl Barez, Martin Tutek +2

How does the conversational past of large language models (LLMs) influence their future performance? Recent work suggests that LLMs are affected by their conversational history in…

cs.CL2025

Distinguishing Ignorance from Error in LLM Hallucinations

Adi Simhi, Jonathan Herzig, Idan Szpektor +1

Large language models (LLMs) are susceptible to hallucinations -- factually incorrect outputs -- leading to a large body of work on detecting and mitigating such cases. We argue th…

cs.CL2024

Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs

Adi Simhi, Jonathan Herzig, Idan Szpektor +1

Large language models (LLMs) are prone to hallucinations, which sparked a widespread effort to detect and prevent them. Recent work attempts to mitigate hallucinations by interveni…