1 paper
Bonan Zhang, Shiyu Dong, Quan Hung Tran +9
Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and i…