14 citations · 18 across the 8 of their papers we have counts for
3 papers · 1 filter
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
Dezhan Tu, Danylo Vashchilenko, Yuzhe Lu +1
Vision-Language Models (VLMs) have demonstrated impressive performance across a versatile set of tasks. A key challenge in accelerating VLMs is storing and accessing the large Key-…
Effectively Fine-tune to Improve Large Multimodal Models for Radiology Report Generation
Yuzhe Lu, Sungmin Hong, Yash Shah +1
Writing radiology reports from medical images requires a high level of domain expertise. It is time-consuming even for trained radiologists and can be error-prone for inexperienced…
On-the-fly Object Detection using StyleGAN with CLIP Guidance
Yuzhe Lu, Shusen Liu, Jayaraman J. Thiagarajan +2
We present a fully automated framework for building object detectors on satellite imagery without requiring any human annotation or intervention. We achieve this by leveraging the…