3 papers
cs.CL2026
VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
Shahrzad Esmat, Dhawal Shah, Ali Jannesari
The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token…
cs.CV2026
LLM-Guided ANN Index Optimization for Human-Object Interaction Retrieval
Shahrzad Esmat, Chaunte W. Lacewell, Sameh Gobriel +2
Retrieval systems underpin modern AI applications -- spanning visual search, recommendation engines, and multi-modal question answering. Modern multi-stage retrieval systems requir…
cs.CV2026
AgenticPruner: MAC-Constrained Neural Network Compression via LLM-Driven Strategy Search
Shahrzad Esmat, Mahdi Banisharif, Ali Jannesari
Neural network pruning remains essential for deploying deep learning models on resource-constrained devices, yet existing approaches primarily target parameter reduction without di…