21 citations · 21 across the 2 of their papers we have counts for
3 papers
cs.CV2025
3rd Place Solution to ICCV LargeFineFoodAI Retrieval
Yang Zhong, Zhiming Wang, Zhaoyang Li +2
This paper introduces the 3rd place solution to the ICCV LargeFineFoodAI Retrieval Competition on Kaggle. Four basic models are independently trained with the weighted sum of ArcFa…
cs.CV2023★ 21 cited
Recognize Anything: A Strong Image Tagging Model
Youcai Zhang, Xinyu Huang, Jinyu Ma +9
We present the Recognize Anything Model (RAM): a strong foundation model for image tagging. RAM makes a substantial step for large models in computer vision, demonstrating the zero…
cs.CV2023
Tag2Text: Guiding Vision-Language Model via Image Tagging
Xinyu Huang, Youcai Zhang, Jinyu Ma +6
This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic…