4 papers · 1 filter
ARMADA: Attribute-Based Multimodal Data Augmentation
Xiaomeng Jin, Jeonghwan Kim, Yu Zhou +4
In Multimodal Language Models (MLMs), the cost of manually annotating high-quality image-text pair data for fine-tuning and alignment is extremely high. While existing multimodal d…
CaLM: Contrasting Large and Small Language Models to Verify Grounded Generation
I-Hung Hsu, Zifeng Wang, Long T. Le +4
Grounded generation aims to equip language models (LMs) with the ability to produce more credible and accountable responses by accurately citing verifiable sources. However, existi…
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
Hritik Bansal, Po-Nien Kung, P. Jeffrey Brantingham +2
Multimodal event argument role labeling (EARL), a task that assigns a role for each event participant (object) in an image is a complex challenge. It requires reasoning over the en…
Argument-Aware Approach To Event Linking
I-Hung Hsu, Zihan Xue, Nilay Pochh +4
Event linking connects event mentions in text with relevant nodes in a knowledge base (KB). Prior research in event linking has mainly borrowed methods from entity linking, overloo…