2 papers
cs.DC2025
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
Shashwat Jaiswal, Shrikara Arun, Anjaly Parayil +8
Low-Rank Adaptation (LoRA) has become the de facto method for parameter-efficient fine-tuning of large language models (LLMs), enabling rapid adaptation to diverse domains. In prod…
cs.CL2025
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
Camille Couturier, Spyros Mastorakis, Haiying Shen +2
Large Language Models (LLMs) are increasingly deployed across edge and cloud platforms for real-time question-answering and retrieval-augmented generation. However, processing leng…