30 citations · 49 across the 6 of their papers we have counts for
20 papers · 1 filter
Bring Metric Functions into Diffusion Models
Jie An, Zhengyuan Yang, Jianfeng Wang +4
We introduce a Cascaded Diffusion Model (Cas-DM) that improves a Denoising Diffusion Probabilistic Model (DDPM) by effectively incorporating additional metric functions in training…
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin +5
In the evolution of Vision-Language Pre-training, shifting from short-text comprehension to encompassing extended textual contexts is pivotal. Recent autoregressive vision-language…
InfoVisDial: An Informative Visual Dialogue Dataset by Bridging Large Multimodal and Language Models
Bingbing Wen, Zhengyuan Yang, Jianfeng Wang +3
In this paper, we build a visual dialogue dataset, named InfoVisDial, which provides rich informative answers in each round even with external knowledge related to the visual conte…
Interfacing Foundation Models' Embeddings
Xueyan Zou, Linjie Li, Jianfeng Wang +10
Foundation models possess strong capabilities in reasoning and memorizing across modalities. To further unleash the power of foundation models, we present FIND, a generalized inter…
Segment and Caption Anything
Xiaoke Huang, Jianfeng Wang, Yansong Tang +5
We propose a method to efficiently equip the Segment Anything Model (SAM) with the ability to generate regional captions. SAM presents strong generalizability to segment anything w…
MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning
Chaoyi Zhang, Kevin Lin, Zhengyuan Yang +5
We present MM-Narrator, a novel system leveraging GPT-4 with multimodal in-context learning for the generation of audio descriptions (AD). Unlike previous methods that primarily fo…