34 citations · 117 across the 19 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024★ 1 cited
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Tianyu Zhang, Suyuchen Wang, Lu Li +6
We introduce Visual Caption Restoration (VCR), a novel vision-language task that challenges models to accurately restore partially obscured texts using pixel-level hints within ima…
cs.CV2024
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
Weiming Ren, Huan Yang, Ge Zhang +4
Image-to-video (I2V) generation aims to use the initial frame (alongside a text prompt) to create a video sequence. A grand challenge in I2V generation is to maintain visual consis…
cs.CV2023
UniIR: Training and Benchmarking Universal Multimodal Information Retrievers
Cong Wei, Yang Chen, Haonan Chen +5
Existing information retrieval (IR) models often assume a homogeneous format, limiting their applicability to diverse user needs, such as searching for images with text description…