2 papers
cs.LG2025
LiteVLM: A Low-Latency Vision-Language Model Inference Pipeline for Resource-Constrained Environments
Jin Huang, Yuchao Jin, Le An +1
This paper introduces an efficient Vision-Language Model (VLM) pipeline specifically optimized for deployment on embedded devices, such as those used in robotics and autonomous dri…
cs.CV2024
ReduceFormer: Attention with Tensor Reduction by Summation
John Yang, Le An, Su Inn Park
Transformers have excelled in many tasks including vision. However, efficient deployment of transformer models in low-latency or high-throughput applications is hindered by the com…