3 papers
cs.LG2026
Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUs
Saipraveen Vabbilisetty, Ajay Kumar Boddepalli, Deep Narayan Mishra +3
Deploying retrieval-augmented generation (RAG) on commodity GPUs such as the NVIDIA T4 (16 GB VRAM) exposes a practical failure mode we call the Compression Paradox: neural prompt…
cs.DC2026
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data
Shashank Kapadia, Deep Narayan Mishra, Sujal Reddy Alugubelli +3
We present SURGE, a streaming GPU encoding system deployed in production to generate embeddings for over 800 million texts across 40,000 logical partitions. Production embedding pi…
cs.LG2026
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
Shashank Kapadia, Deep Naryan Mishra, Sujal Reddy Alugubelli +4
Layer-aligned distillation and convergence-based early exit represent two predominant computational efficiency paradigms for transformer inference; yet we establish that they exhib…