activity
20222024
most citedMiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

67 citations · 129 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta +4

Leveraging Large Language Models' remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. Howeve…

cs.CV202367 cited

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Jun Chen, Deyao Zhu, Xiaoqian Shen +7

Large language models have shown their remarkable capabilities as a general interface for various language-related applications. Motivated by this, we target to build a unified int…

cs.CV20234 cited

Exploring Open-Vocabulary Semantic Segmentation without Human Labels

Jun Chen, Deyao Zhu, Guocheng Qian +6

Semantic segmentation is a crucial task in computer vision that involves segmenting images into semantically meaningful regions at the pixel level. However, existing approaches oft…

cs.CV202340 cited

ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Deyao Zhu, Jun Chen, Kilichbek Haydarov +3

Asking insightful questions is crucial for acquiring knowledge and expanding our understanding of the world. However, the importance of questioning has been largely overlooked in A…

cs.CV20228 cited

Few-Shot Class-Incremental Learning via Entropy-Regularized Data-Free Replay

Huan Liu, Li Gu, Zhixiang Chi +4

Few-shot class-incremental learning (FSCIL) has been proposed aiming to enable a deep learning system to incrementally learn new classes with limited data. Recently, a pioneer clai…