1 paper
Xiaoyu Liang, Chaofeng Guan, Jiaying Lu +3
Vision-Language Models (VLMs) have achieved notable success in multimodal tasks but face practical limitations due to the quadratic complexity of decoder attention mechanisms and a…