most citedLLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

119 citations · 210 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20231 cited

Towards Unified and Effective Domain Generalization

Yiyuan Zhang, Kaixiong Gong, Xiaohan Ding +4

We propose , a novel and fied framework for omain eneralization that is capable of significantly enhancing the out-of-distribu…

cs.MM202325 cited

ImageBind-LLM: Multi-modality Instruction Tuning

Jiaming Han, Renrui Zhang, Wenqi Shao +14

We present ImageBind-LLM, a multi-modality instruction tuning method of large language models (LLMs) via ImageBind. Existing works mainly focus on language and image instruction tu…

cs.CV202347 cited

Meta-Transformer: A Unified Framework for Multimodal Learning

Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang +4

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging…

cs.CV20232 cited

Space Engage: Collaborative Space Supervision for Contrastive-based Semi-Supervised Semantic Segmentation

Changqi Wang, Haoyu Xie, Yuhui Yuan +2

Semi-Supervised Semantic Segmentation (S4) aims to train a segmentation model with limited labeled images and a substantial volume of unlabeled images. To improve the robustness of…

cs.CV2023119 cited

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Peng Gao, Jiaming Han, Renrui Zhang +9

How to efficiently transform large language models (LLMs) into instruction followers is recently a popular research direction, while training LLM for multi-modal reasoning remains…

cs.CV202216 cited

Prompt Vision Transformer for Domain Generalization

Zangwei Zheng, Xiangyu Yue, Kai Wang +1

Though vision transformers (ViTs) have exhibited impressive ability for representation learning, we empirically find that they cannot generalize well to unseen domains with previou…