2 papers
cs.AI2026
GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference
Vimal William, Ravi Tandon, Jyotikrishna Dass
As Large Language Models scale to increasingly long contexts, the memory I/O and computational overhead of the Key-Value (KV) cache during decoding emerges as the primary throughpu…
cs.AR2025
TYTAN: Taylor-series based Non-Linear Activation Engine for Deep Learning Accelerators
Soham Pramanik, Vimal William, Arnab Raha +3
The rapid advancement in AI architectures and the proliferation of AI-enabled systems have intensified the need for domain-specific architectures that enhance both the acceleration…