2 papers
cs.LG2025
VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
Kichang Yang, Seonjun Kim, Minjae Kim +3
Edge deployment of large Vision-Language Models (VLMs) increasingly relies on flash-based weight offloading, where activation sparsification is used to reduce I/O overhead. However…
cs.CV2024
Mondrian: On-Device High-Performance Video Analytics with Compressive Packed Inference
Changmin Jeon, Seonjun Kim, Juheon Yi +1
In this paper, we present Mondrian, an edge system that enables high-performance object detection on high-resolution video streams. Many lightweight models and system optimization…