27 citations · 33 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 27 cited
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Feng Li, Renrui Zhang, Hao Zhang +5
Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image t…
cs.CV2024
Perceptive self-supervised learning network for noisy image watermark removal
Chunwei Tian, Menghua Zheng, Bo Li +3
Popular methods usually use a degradation model in a supervised way to learn a watermark removal model. However, it is true that reference images are difficult to obtain in the rea…
cs.CV2023★ 6 cited
OtterHD: A High-Resolution Multi-modality Model
Bo Li, Peiyuan Zhang, Jingkang Yang +3
In this paper, we present OtterHD-8B, an innovative multimodal model evolved from Fuyu-8B, specifically engineered to interpret high-resolution visual inputs with granular precisio…