1 paper
Kichang Yang, Seonjun Kim, Minjae Kim +3
Edge deployment of large Vision-Language Models (VLMs) increasingly relies on flash-based weight offloading, where activation sparsification is used to reduce I/O overhead. However…