activity
20162023
most citedMIMIC-IT: Multi-Modal In-Context Instruction Tuning

28 citations · 55 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV20236 cited

OtterHD: A High-Resolution Multi-modality Model

Bo Li, Peiyuan Zhang, Jingkang Yang +3

In this paper, we present OtterHD-8B, an innovative multimodal model evolved from Fuyu-8B, specifically engineered to interpret high-resolution visual inputs with granular precisio…

cs.CV202314 cited

Large Language Models are Visual Reasoning Coordinators

Liangyu Chen, Bo Li, Sheng Shen +5

Visual reasoning requires multimodal perception and commonsense cognition of the world. Recently, multiple vision-language models (VLMs) have been proposed with excellent commonsen…

cs.CV202328 cited

MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Bo Li, Yuanhan Zhang, Liangyu Chen +5

High-quality instructions and responses are essential for the zero-shot performance of large language models on interactive natural language tasks. For interactive vision-language…

cs.CV2023

SAD: Segment Any RGBD

Jun Cen, Yizheng Wu, Kewei Wang +6

The Segment Anything Model (SAM) has demonstrated its effectiveness in segmenting any part of 2D RGB images. However, SAM exhibits a stronger emphasis on texture information while…

cs.CV20221 cited

On-Device Domain Generalization

Kaiyang Zhou, Yuanhan Zhang, Yuhang Zang +3

We present a systematic study of domain generalization (DG) for tiny neural networks. This problem is critical to on-device machine learning applications but has been overlooked in…

cs.CV2022

Panoptic Scene Graph Generation

Jingkang Yang, Yi Zhe Ang, Zujin Guo +3

Existing research addresses scene graph generation (SGG) -- a critical technology for scene understanding in images -- from a detection perspective, i.e., objects are detected usin…