179 citations · 199 across the 19 of their papers we have counts for
Showing cs.ARShow all
2 papers · 1 filter
cs.AR2025
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
Hang Zhang, Jiuchen Shi, Yixiao Wang +3
Multiple Low-Rank Adapters (Multi-LoRAs) are gaining popularity for task-specific Large Language Model (LLM) applications. For multi-LoRA serving, caching hot KV caches and LoRA ad…
cs.AR2024
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
Weichuang Zhang, Jieru Zhao, Guan Shen +3
With the increasing demand for computing capability given limited resource and power budgets, it is crucial to deploy applications to customized accelerators like FPGAs. However, F…