3 papers
cs.NI2025
Inferring Causal Relationships to Improve Caching for Clients with Correlated Requests: Applications to VR
Agrim Bari, Gustavo de Veciana, Yuqi Zhou
Efficient edge caching reduces latency and alleviates backhaul congestion in modern networks. Traditional caching policies, such as Least Recently Used (LRU) and Least Frequently U…
cs.LG2025
Optimal Scheduling Algorithms for LLM Inference: Theory and Practice
Agrim Bari, Parikshit Hegde, Gustavo de Veciana
With the growing use of Large Language Model (LLM)-based tools like ChatGPT, Perplexity, and Gemini across industries, there is a rising need for efficient LLM inference systems. T…
cs.NI2025
Fundamentals of Caching Layered Data objects
Agrim Bari, Gustavo de Veciana, George Kesidis
The effective management of large amounts of data processed or required by today's cloud or edge computing systems remains a fundamental challenge. This paper focuses on cache mana…