3 papers
cs.AI2026
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference
Srihari Unnikrishnan
Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request. Existing prefix-caching systems reduce…
cs.CL2026
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
Srihari Unnikrishnan, Jaskaran Singh Walia, Drishti Goel +1
Cloud services experience frequent incidents that require rapid diagnosis and resolution. Troubleshooting guides help engineers respond consistently, but creating them manually is…
q-fin.CP2026
Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation
Jaskaran Singh Walia, Aarush Sinha, Naman Saraswat +2
Financial bond yield forecasting is challenging due to data scarcity, nonlinear macroeconomic dependencies, and evolving market conditions. In this paper, we propose a novel framew…