activity
20212024
most citedMuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and Grounding

1 citations · 4 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20241 cited

Rethinking Interactive Image Segmentation with Low Latency, High Quality, and Diverse Prompts

Qin Liu, Jaemin Cho, Mohit Bansal +1

The goal of interactive image segmentation is to delineate specific regions within an image via visual or language prompts. Low-latency and high-quality interactive segmentation wi…

cs.CV2024

SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data

Jialu Li, Jaemin Cho, Yi-Lin Sung +2

Recent text-to-image (T2I) generation models have demonstrated impressive capabilities in creating images from text descriptions. However, these T2I generation models often fall sh…

cs.CV20241 cited

Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training

David Wan, Jaemin Cho, Elias Stengel-Eskin +1

Highlighting particularly relevant regions of an image can improve the performance of vision-language models (VLMs) on various vision-language (VL) tasks by guiding the model to at…

cs.CV20231 cited

Hierarchical Video-Moment Retrieval and Step-Captioning

Abhay Zala, Jaemin Cho, Satwik Kottur +4

There is growing interest in searching for information from large video corpora. Prior works have studied relevant tasks, such as text-based video retrieval, moment retrieval, vide…

cs.CL20211 cited

MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and Grounding

Revanth Gangi Reddy, Xilin Rui, Manling Li +9

Recently, there has been an increasing interest in building question answering (QA) models that reason across multiple modalities, such as text and images. However, QA using images…