12 papers
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
Nitay Calderon, Eyal Ben-David, Zorik Gekhman +2
Standard factuality evaluations of LLMs treat all errors alike, obscuring whether failures arise from missing knowledge (empty shelves) or from limited access to encoded facts (los…
LLM Explainability with Counterfactual Chains and Causal Graphs
Nirit Nussbaum-Hoffer, Nitay Calderon, Liat Ein-Dor +1
Causal graphs provide a high-level language for making mechanisms transparent. Recent work uses Large Language Models (LLMs) to recover causal graphs of external-world processes. I…
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Tomer Keren, Nitay Calderon, Asaf Yehudai +3
As agent capabilities advance, existing benchmarks, such as -Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and lab…
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
Gilat Toker, Nitay Calderon, Ohad Amosy +1
Concept-based explanations quantify how high-level concepts (e.g., gender or experience) influence model behavior, which is crucial for decision-makers in high-stakes domains. Rece…
Leveraging NTPs for Efficient Hallucination Detection in VLMs
Ofir Azachi, Kfir Eliyahu, Eyal El Ani +4
Hallucinations of vision-language models (VLMs), which are misalignments between visual content and generated text, undermine the reliability of VLMs. One common approach for detec…
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance
Omer Nahum, Nitay Calderon, Orgad Keller +2
NLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field. Traditionally, expert annotations ensure high-quality label…