Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
Dezhan Tu, Danylo Vashchilenko, Yuzhe Lu +1
Vision-Language Models (VLMs) have demonstrated impressive performance across a versatile set of tasks. A key challenge in accelerating VLMs is storing and accessing the large Key-…
cs.CV2024
Customize Your Own Paired Data via Few-shot Way
Jinshu Chen, Bingchuan Li, Miao Hua +2
Existing solutions to image editing tasks suffer from several issues. Though achieving remarkably satisfying generated results, some supervised methods require huge amounts of pair…
cs.CV2023
Effectively Fine-tune to Improve Large Multimodal Models for Radiology Report Generation
Yuzhe Lu, Sungmin Hong, Yash Shah +1
Writing radiology reports from medical images requires a high level of domain expertise. It is time-consuming even for trained radiologists and can be error-prone for inexperienced…