1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.DC2025★ 1 cited
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
Shubham Agarwal, Sai Sundaresan, Subrata Mitra +6
Retrieval-Augmented Generation (RAG) is often used with Large Language Models (LLMs) to infuse domain knowledge or user-specific information. In RAG, given a user query, a retrieve…
cs.LG2025
Prompt-Aware Scheduling for Efficient Text-to-Image Inferencing System
Shubham Agarwal, Saud Iqbal, Subrata Mitra
Traditional ML models utilize controlled approximations during high loads, employing faster, but less accurate models in a process called accuracy scaling. However, this method is…