2 papers
cs.LG2025
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
Sai Gokhale, Devleena Das, Rajeev Patwari +2
Long-context Large Language Models (LLMs) face significant memory bottlenecks during inference due to the linear growth of key-value (KV) cache with sequence length. While individu…
quant-ph2025
Generation and Detection of Hyperentangled Bell States at an Ultra-High Flux
Netanel P. Yaish, Samata Gokhale, Avi Peer
We demonstrate both the generation and detection of an ultra-high flux of polarization Bell states using broadband hyper-entangled bi-photons that are quantum-correlated in both po…