125 citations · 349 across the 13 of their papers we have counts for
4 papers · 1 filter
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
Tianxiang Hao, Mengyao Lyu, Hui Chen +4
With the advancement of large pre-trained vision-language models, effectively transferring the knowledge embedded within these foundational models to downstream tasks has become a…
Evolving Semantic Prototype Improves Generative Zero-Shot Learning
Shiming Chen, Wenjin Hou, Ziming Hong +5
In zero-shot learning (ZSL), generative methods synthesize class-related sample features based on predefined semantic prototypes. They advance the ZSL performance by synthesizing u…
Sticker820K: Empowering Interactive Retrieval with Stickers
Sijie Zhao, Yixiao Ge, Zhongang Qi +4
Stickers have become a ubiquitous part of modern-day communication, conveying complex emotions through visual imagery. To facilitate the development of more powerful algorithms for…
What Makes for Good Visual Tokenizers for Large Language Models?
Guangzhi Wang, Yixiao Ge, Xiaohan Ding +2
We empirically investigate proper pre-training methods to build good visual tokenizers, making Large Language Models (LLMs) powerful Multimodal Large Language Models (MLLMs). In ou…