3 citations · 3 across the 1 of their papers we have counts for
3 papers
cs.CV2024
TokenPacker: Efficient Visual Projector for Multimodal LLM
Wentong Li, Yuqian Yuan, Jian Liu +5
The visual projector serves as an essential bridge between the visual encoder and the Large Language Model (LLM) in a Multimodal LLM (MLLM). Typically, MLLMs adopt a simple MLP to…
cs.CV2024
Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label Classification
Xuelin Zhu, Jian Liu, Dongqi Tang +4
Identifying labels that did not appear during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. To this end, recent studies have attempte…
cs.CV2023★ 3 cited
Label-efficient Segmentation via Affinity Propagation
Wentong Li, Yuqian Yuan, Song Wang +5
Weakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, whil…