22 papers · 1 filter
Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
Eran Hirsch, David Wan, Han Wang +3
Deep research (DR) systems produce long-form cited reports by orchestrating multiple agents that search and synthesize information from the web. Citations are the primary mechanism…
Effective QA-driven Annotation of Predicate-Argument Relations Across Languages
Jonathan Davidov, Aviv Slobodkin, Shmuel Tomi Klein +3
Explicit representations of predicate-argument relations form the basis of interpretable semantic analysis, supporting reasoning, generation, and evaluation. However, attaining suc…
User-Centric Evidence Ranking for Attribution and Fact Verification
Guy Alt, Eran Hirsch, Serwar Basch +2
Attribution and fact verification are critical challenges in natural language processing for assessing information reliability. While automated systems and Large Language Models (L…
QA-Noun: Representing Nominal Semantics via Natural Language Question-Answer Pairs
Maria Tseytlin, Paul Roit, Omri Abend +2
Decomposing sentences into fine-grained meaning units is increasingly used to model semantic alignment. While QA-based semantic approaches have shown effectiveness for representing…
PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
Sapir Harary, Eran Hirsch, Aviv Slobodkin +3
Natural Language Inference (NLI) models have been used in various ways to improve the factuality of LLM outputs. This is typically done by applying an NLI model to judge whether th…
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
Shiyue Zhang, David Wan, Arie Cattan +3
How to properly conduct human evaluations for text summarization is a longstanding challenge. The Pyramid human evaluation protocol, which assesses content selection by breaking th…