1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2024
Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models
Songtao Jiang, Yan Zhang, Chenyi Zhou +4
Multimodal Large Language Models (MLLMs) such as GPT-4V and Gemini Pro face challenges in achieving human-level perception in Visual Question Answering (VQA), particularly in objec…
cs.CV2023★ 1 cited
USER: Unified Semantic Enhancement with Momentum Contrast for Image-Text Retrieval
Yan Zhang, Zhong Ji, Di Wang +2
As a fundamental and challenging task in bridging language and vision domains, Image-Text Retrieval (ITR) aims at searching for the target instances that are semantically relevant…
cs.LG2021
Learning to Represent and Predict Sets with Deep Neural Networks
Yan Zhang
In this thesis, we develop various techniques for working with sets in machine learning. Each input or output is not an image or a sequence, but a set: an unordered collection of m…