1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
LinVT: Empower Your Image-level Large Language Model to Understand Videos
Lishuai Gao, Yujie Zhong, Yingsen Zeng +3
Large Language Models (LLMs) have been widely used in various tasks, motivating us to develop an LLM-based assistant for videos. Instead of training from scratch, we propose a modu…
cs.CV2024
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
Cong Wei, Yujie Zhong, Haoxian Tan +4
This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite signif…
cs.CV2024★ 1 cited
LaSagnA: Language-based Segmentation Assistant for Complex Queries
Cong Wei, Haoxian Tan, Yujie Zhong +2
Recent advancements have empowered Large Language Models for Vision (vLLMs) to generate detailed perceptual outcomes, including bounding boxes and masks. Nonetheless, there are two…