1 paper
Paribesh Regmi, Qingshuang Chen, Chi Zhang +3
Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deploy…