most citedLabel-efficient Segmentation via Affinity Propagation

3 citations · 3 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV2024

Scalable Autoregressive Monocular Depth Estimation

Jinhong Wang, Jian Liu, Dongqi Tang +5

This paper shows that the autoregressive model is an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with…

cs.CV2024

TokenPacker: Efficient Visual Projector for Multimodal LLM

Wentong Li, Yuqian Yuan, Jian Liu +5

The visual projector serves as an essential bridge between the visual encoder and the Large Language Model (LLM) in a Multimodal LLM (MLLM). Typically, MLLMs adopt a simple MLP to…

cs.CV2024

Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label Classification

Xuelin Zhu, Jian Liu, Dongqi Tang +4

Identifying labels that did not appear during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. To this end, recent studies have attempte…

cs.CV2023

Text as Image: Learning Transferable Adapter for Multi-Label Classification

Xuelin Zhu, Jiuxin Cao, Jian liu +7

Pre-trained vision-language models have notably accelerated progress of open-world concept recognition. Their impressive zero-shot ability has recently been transferred to multi-la…

cs.CV2023

Osprey: Pixel Understanding with Visual Instruction Tuning

Yuqian Yuan, Wentong Li, Jian Liu +5

Multimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However, current MLLMs pr…

cs.CV20233 cited

Label-efficient Segmentation via Affinity Propagation

Wentong Li, Yuqian Yuan, Song Wang +5

Weakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, whil…