10 papers · 1 filter
Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
Eran Hirsch, David Wan, Han Wang +3
Deep research (DR) systems produce long-form cited reports by orchestrating multiple agents that search and synthesize information from the web. Citations are the primary mechanism…
User-Centric Evidence Ranking for Attribution and Fact Verification
Guy Alt, Eran Hirsch, Serwar Basch +2
Attribution and fact verification are critical challenges in natural language processing for assessing information reliability. While automated systems and Large Language Models (L…
PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
Sapir Harary, Eran Hirsch, Aviv Slobodkin +3
Natural Language Inference (NLI) models have been used in various ways to improve the factuality of LLM outputs. This is typically done by applying an NLI model to judge whether th…
CRISP: Complex Reasoning with Interpretable Step-based Plans
Matan Vetzler, Koren Lazar, Guy Uziel +3
Recent advancements in large language models (LLMs) underscore the need for stronger reasoning capabilities to solve complex problems effectively. While Chain-of-Thought (CoT) reas…
GenerationPrograms: Fine-grained Attribution with Executable Programs
David Wan, Eran Hirsch, Elias Stengel-Eskin +2
Recent large language models (LLMs) achieve impressive performance in source-conditioned text generation but often fail to correctly provide fine-grained attributions for their out…
CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection
Ron Eliav, Arie Cattan, Eran Hirsch +4
A common approach to hallucination detection casts it as a natural language inference (NLI) task, often using LLMs to classify whether the generated text is entailed by correspondi…