papers

Publications (20)

cs.CV2024

Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training

David Wan, Jaemin Cho, Elias Stengel-Eskin +1

Highlighting particularly relevant regions of an image can improve the performance of vision-language models (VLMs) on various vision-language (VL) tasks by guiding the model to at…

cs.CL2025

GenerationPrograms: Fine-grained Attribution with Executable Programs

David Wan, Eran Hirsch, Elias Stengel-Eskin +2

Recent large language models (LLMs) achieve impressive performance in source-conditioned text generation but often fail to correctly provide fine-grained attributions for their out…

cs.CL2023

Faithfulness-Aware Decoding Strategies for Abstractive Summarization

David Wan, Mengwen Liu, Kathleen McKeown +2

Despite significant progress in understanding and improving faithfulness in abstractive summarization, the question of how decoding strategies affect faithfulness is less studied.…

cs.CL2026

Multimodal Fact-Level Attribution for Verifiable Reasoning

David Wan, Han Wang, Ziyang Wang +3

Multimodal large language models (MLLMs) are increasingly used for real-world tasks involving multi-step reasoning and long-form generation, where reliability requires grounding mo…

cs.CL2025

QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization

Shiyue Zhang, David Wan, Arie Cattan +3

How to properly conduct human evaluations for text summarization is a longstanding challenge. The Pyramid human evaluation protocol, which assesses content selection by breaking th…

cs.CL2021

Segmenting Subtitles for Correcting ASR Segmentation Errors

David Wan, Chris Kedzie, Faisal Ladhak +6

Typical ASR systems segment the input audio into utterances using purely acoustic information, which may not resemble the sentence-like units that are expected by conventional mach…

cs.CL2020

Subtitles to Segmentation: Improving Low-Resource Speech-to-Text Translation Pipelines

David Wan, Zhengping Jiang, Chris Kedzie +3

In this work, we focus on improving ASR output segmentation in the context of low-resource language speech-to-text translation. ASR output segmentation is crucial, as ASR systems s…

cs.CL2022

FactPEGASUS: Factuality-Aware Pre-training and Fine-tuning for Abstractive Summarization

David Wan, Mohit Bansal

We present FactPEGASUS, an abstractive summarization model that addresses the problem of factuality during pre-training and fine-tuning: (1) We augment the sentence selection strat…

cs.CL2025

PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise

Sapir Harary, Eran Hirsch, Aviv Slobodkin +3

Natural Language Inference (NLI) models have been used in various ways to improve the factuality of LLM outputs. This is typically done by applying an NLI model to judge whether th…

cs.CL2023

HistAlign: Improving Context Dependency in Language Generation by Aligning with History

David Wan, Shiyue Zhang, Mohit Bansal

Language models (LMs) can generate hallucinations and incoherent outputs, which highlights their weak context dependency. Cache-LMs, which augment LMs with a memory of recent histo…

cs.CL2025

On Positional Bias of Faithfulness for Long-form Summarization

David Wan, Jesse Vig, Mohit Bansal +1

Large Language Models (LLMs) often exhibit positional bias in long-context settings, under-attending to information in the middle of inputs. We investigate the presence of this bia…

cs.CL2025

LAQuer: Localized Attribution Queries in Content-grounded Generation

Eran Hirsch, Aviv Slobodkin, David Wan +3

Grounded text generation models often produce content that deviates from their source material, requiring user verification to ensure accuracy. Existing attribution methods associa…

cs.CL2026

MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments

Han Wang, David Wan, Hyunji Lee +6

Motivated by the underspecified, multi-hop nature of search queries and the multimodal, heterogeneous, and often conflicting nature of real-world web results, we introduce MERRIN (…

cs.CV2025

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

David Wan, Han Wang, Elias Stengel-Eskin +2

Online video web content is richly multimodal: a single video blends vision, speech, ambient audio, and on-screen text. Retrieval systems typically treat these modalities as indepe…

cs.CL2022

Evaluating and Improving Factuality in Multimodal Abstractive Summarization

David Wan, Mohit Bansal

Current metrics for evaluating factuality for abstractive document summarization have achieved high correlations with human judgment, but they do not account for the vision modalit…

cs.CL2023

Extractive is not Faithful: An Investigation of Broad Unfaithfulness Problems in Extractive Summarization

Shiyue Zhang, David Wan, Mohit Bansal

The problems of unfaithful summaries have been widely discussed under the context of abstractive summarization. Though extractive summarization is less prone to the common unfaithf…

cs.CL2025

Localizing Factual Inconsistencies in Attributable Text Generation

Arie Cattan, Paul Roit, Shiyue Zhang +5

There has been an increasing interest in detecting hallucinations in model-generated texts, both manually and automatically, at varying levels of granularity. However, most existin…

cs.CL2025

MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration

David Wan, Justin Chih-Yao Chen, Elias Stengel-Eskin +1

Multi-agent collaboration among models has shown promise in reasoning tasks but is underexplored in long-form generation tasks like summarization and question-answering. We extend…

cs.CL2025

DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning

Nithin Sivakumaran, Justin Chih-Yao Chen, David Wan +4

Specialized visual tools can augment large language models or vision language models with expert knowledge (e.g., grounding, spatial reasoning, medical knowledge, etc.), but knowin…

cs.CL2020

Incorporating Terminology Constraints in Automatic Post-Editing

David Wan, Chris Kedzie, Faisal Ladhak +2

Users of machine translation (MT) may want to ensure the use of specific lexical terminologies. While there exist techniques for incorporating terminology constraints during infere…