activity
20162024
most citedVariational Autoencoder for Deep Learning of Images, Labels and Captions

371 citations · 1k across the 29 of their papers we have counts for

collaborators

17 papers

cs.CV202328 cited

MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Bo Li, Yuanhan Zhang, Liangyu Chen +5

High-quality instructions and responses are essential for the zero-shot performance of large language models on interactive natural language tasks. For interactive vision-language…

cs.CV2023231 cited

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Chunyuan Li, Cliff Wong, Sheng Zhang +6

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversation…

cs.CL2023189 cited

Instruction Tuning with GPT-4

Baolin Peng, Chunyuan Li, Pengcheng He +2

Prior work has shown that finetuning large language models (LLMs) using machine-generated instruction-following data enables such models to achieve remarkable zero-shot capabilitie…

cs.CV20232 cited

A Simple Framework for Open-Vocabulary Segmentation and Detection

Hao Zhang, Feng Li, Xueyan Zou +5

We present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets. To bridge the gap of voca…

cs.CV20231 cited

Scaling Vision-Language Models with Sparse Mixture of Experts

Sheng Shen, Zhewei Yao, Chunyuan Li +3

The field of natural language processing (NLP) has made significant strides in recent years, particularly in the development of large-scale vision-language models (VLMs). These mod…

cs.CV2023

Learning Customized Visual Models with Retrieval-Augmented Knowledge

Haotian Liu, Kilho Son, Jianwei Yang +4

Image-text contrastive learning models such as CLIP have demonstrated strong task transfer ability. The high generality and usability of these visual models is achieved via a web-s…