4 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2024
EVLM: An Efficient Vision-Language Model for Visual Understanding
Kaibing Chen, Dong Shen, Hanwen Zhong +14
In the field of multi-modal language models, the majority of methods are built on an architecture similar to LLaVA. These models use a single-layer ViT feature as a visual prompt,…
cs.MM2023
Parameter-Efficient Transfer Learning for Audio-Visual-Language Tasks
Hongye Liu, Xianhai Xie, Yang Gao +2
The pretrain-then-finetune paradigm has been widely used in various unimodal and multimodal tasks. However, finetuning all the parameters of a pre-trained model becomes prohibitive…
cs.CV2022★ 4 cited
Real-time End-to-End Video Text Spotter with Contrastive Representation Learning
Wejia Wu, Zhuang Li, Jiahong Li +5
Video text spotting(VTS) is the task that requires simultaneously detecting, tracking and recognizing text in the video. Existing video text spotting methods typically develop soph…