2 papers
cs.DC2026
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
Satyam Kumar, Arpit Singh Gautam, Kailash Talreja +1
Efficient LLM serving must balance throughput and latency across diverse, bursty workloads. We introduce StreamServe, a disaggregated prefill decode serving architecture that combi…
cs.CL2026
The Energy of Falsehood: Detecting Hallucinations via Diffusion Model Likelihoods
Arpit Singh Gautam, Kailash Talreja, Saurabh Jha
Large Language Models (LLMs) frequently hallucinate plausible but incorrect assertions, a vulnerability often missed by uncertainty metrics when models are confidently wrong. We pr…