most citedGRiT: A Generative Region-to-text Transformer for Object Understanding

30 citations · 49 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

20 papers · 1 filter

cs.CV2024

Bring Metric Functions into Diffusion Models

Jie An, Zhengyuan Yang, Jianfeng Wang +4

We introduce a Cascaded Diffusion Model (Cas-DM) that improves a Denoising Diffusion Probabilistic Model (DDPM) by effectively incorporating additional metric functions in training…

cs.CV2024

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin +5

In the evolution of Vision-Language Pre-training, shifting from short-text comprehension to encompassing extended textual contexts is pivotal. Recent autoregressive vision-language…

cs.CV2023

InfoVisDial: An Informative Visual Dialogue Dataset by Bridging Large Multimodal and Language Models

Bingbing Wen, Zhengyuan Yang, Jianfeng Wang +3

In this paper, we build a visual dialogue dataset, named InfoVisDial, which provides rich informative answers in each round even with external knowledge related to the visual conte…

cs.CV2023

Interfacing Foundation Models' Embeddings

Xueyan Zou, Linjie Li, Jianfeng Wang +10

Foundation models possess strong capabilities in reasoning and memorizing across modalities. To further unleash the power of foundation models, we present FIND, a generalized inter…

cs.CV2023

Segment and Caption Anything

Xiaoke Huang, Jianfeng Wang, Yansong Tang +5

We propose a method to efficiently equip the Segment Anything Model (SAM) with the ability to generate regional captions. SAM presents strong generalizability to segment anything w…

cs.CV2023

MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning

Chaoyi Zhang, Kevin Lin, Zhengyuan Yang +5

We present MM-Narrator, a novel system leveraging GPT-4 with multimodal in-context learning for the generation of audio descriptions (AD). Unlike previous methods that primarily fo…