2 papers
cs.AR2026
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
Rakshith Jayanth, Viktor Prasanna
In long-context large language model (LLM) inference, the prefill stage dominates computation due to self-attention over the complete input context. Sparse attention significantly…
cs.DC2025
Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
Neelesh Gupta, Rakshith Jayanth, Dhruv Parikh +1
The proliferation of large language models has driven demand for long-context inference on resource-constrained edge platforms. However, deploying these models on Neural Processing…