activity
20162023
most citedVisual Instruction Tuning

699 citations · 2.3k across the 34 of their papers we have counts for

collaborators
Showing cs.CVShow all

20 papers · 1 filter

cs.CV2023★ 8 cited

Large Multimodal Models: Notes on CVPR 2023 Tutorial

Chunyuan Li

This tutorial note summarizes the presentation on ``Large Multimodal Models: Towards Building and Surpassing Multimodal GPT-4'', a part of CVPR 2023 tutorial on ``Recent Advances i…

cs.CV2023★ 231 cited

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Chunyuan Li, Cliff Wong, Sheng Zhang +6

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversation…

cs.CV2023★ 699 cited

Visual Instruction Tuning

Haotian Liu, Chunyuan Li, Qingyang Wu +1

Instruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored i…

cs.CV2023★ 1 cited

Scaling Vision-Language Models with Sparse Mixture of Experts

Sheng Shen, Zhewei Yao, Chunyuan Li +3

The field of natural language processing (NLP) has made significant strides in recent years, particularly in the development of large-scale vision-language models (VLMs). These mod…

cs.CV2023

Learning Customized Visual Models with Retrieval-Augmented Knowledge

Haotian Liu, Kilho Son, Jianwei Yang +4

Image-text contrastive learning models such as CLIP have demonstrated strong task transfer ability. The high generality and usability of these visual models is achieved via a web-s…

cs.CV2023★ 3 cited

GLIGEN: Open-Set Grounded Text-to-Image Generation

Yuheng Li, Haotian Liu, Qingyang Wu +5

Large-scale text-to-image diffusion models have made amazing advances. However, the status quo is to use text input alone, which can impede controllability. In this work, we propos…