2 papers
cs.LG2026
vCache: Verified Semantic Prompt Caching
Luis Gaspar Schroeder, Aditya Desai, Alejandro Cuadron +7
Semantic caches return cached responses for semantically similar prompts to reduce LLM inference latency and cost. They embed cached prompts and store them alongside their response…
cs.CL2024
Dynamic Fog Computing for Enhanced LLM Execution in Medical Applications
Philipp Zagar, Vishnu Ravi, Lauren Aalami +3
The ability of large language models (LLMs) to transform, interpret, and comprehend vast quantities of heterogeneous data presents a significant opportunity to enhance data-driven…