4 citations · 7 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 3 cited
Open-Set Image Tagging with Multi-Grained Text Supervision
Xinyu Huang, Yi-Jie Huang, Youcai Zhang +6
In this paper, we introduce the Recognize Anything Plus Model (RAM++), an open-set image tagging model effectively leveraging multi-grained text supervision. Previous approaches (e…
cs.CV2023★ 4 cited
u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
Jinjin Xu, Liwu Xu, Yuzhe Yang +5
Recent advancements in multi-modal large language models (MLLMs) have led to substantial improvements in visual understanding, primarily driven by sophisticated modality alignment…