collaborators

5 papers

cs.AI2026

AgentKVShift: Efficient KV Cache Reuse for Agentic Memory Systems

Nilesh Prasad Pandey, Jason Kong, Lanxiang Hu +5

Memory-augmented LLM agents maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content with LLM-generated metadata such…

cs.LG2026

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models

Jason Kong, Nilesh Prasad Pandey, Flavio Ponzina +1

Deploying Large Language Models (LLMs) on edge devices faces severe computational and memory constraints, limiting real-time processing and on-device intelligence. Hybrid architect…

cs.LG2026

QMC: Efficient SLM Edge Inference via Outlier-Aware Quantization and Emergent Memories Co-Design

Nilesh Prasad Pandey, Jangseon Park, Onat Gungor +2

Deploying Small Language Models (SLMs) on edge platforms is critical for real-time, privacy-sensitive generative AI, yet constrained by memory, latency, and energy budgets. Quantiz…

cs.LG2025

DPQ-HD: Post-Training Compression for Ultra-Low Power Hyperdimensional Computing

Nilesh Prasad Pandey, Shriniwas Kulkarni, David Wang +3

Hyperdimensional Computing (HDC) is emerging as a promising approach for edge AI, offering a balance between accuracy and efficiency. However, current HDC-based applications often…

cs.LG2025

Sparse High Rank Adapters

Kartikeya Bhardwaj, Nilesh Prasad Pandey, Sweta Priyadarshi +9

Low Rank Adaptation (LoRA) has gained massive attention in the recent generative AI research. One of the main advantages of LoRA is its ability to be fused with pretrained models,…