1 paper
Saeed Khaki, Nima Safaei, Kamal Ginotra
Transformer-based vision-language models (VLMs) contain substantial depth redundancy, yet the effect of removing specific decoder layers remains poorly understood, especially for d…