most citedMiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

67 citations · 122 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

How Well Can Vision Language Models See Image Details?

Chenhui Gou, Abdulwahab Felemban, Faizan Farooq Khan +4

Large Language Model-based Vision-Language Models (LLM-based VLMs) have demonstrated impressive results in various vision-language understanding tasks. However, how well these VLMs…

cs.CV2024

Goldfish: Vision-Language Understanding of Arbitrarily Long Videos

Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman +6

Most current LLM-based models for video understanding can process videos within minutes. However, they struggle with lengthy videos due to challenges such as "noise and redundancy"…

cs.AI20245 cited

MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis

Asma Alkhaldi, Raneem Alnajim, Layan Alabdullatef +5

Recent advancements in artificial intelligence (AI) have precipitated significant breakthroughs in healthcare, particularly in refining diagnostic procedures. However, previous stu…

cs.CV20246 cited

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman +4

This paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding. The model is capable of processing both temporal visual…

cs.CV202367 cited

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Jun Chen, Deyao Zhu, Xiaoqian Shen +7

Large language models have shown their remarkable capabilities as a general interface for various language-related applications. Motivated by this, we target to build a unified int…

cs.CV20234 cited

Exploring Open-Vocabulary Semantic Segmentation without Human Labels

Jun Chen, Deyao Zhu, Guocheng Qian +6

Semantic segmentation is a crucial task in computer vision that involves segmenting images into semantically meaningful regions at the pixel level. However, existing approaches oft…