3 papers
cs.AR2026
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
Rakshith Jayanth, Viktor Prasanna
In long-context large language model (LLM) inference, the prefill stage dominates computation due to self-attention over the complete input context. Sparse attention significantly…
cs.DC2025
Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
Neelesh Gupta, Rakshith Jayanth, Dhruv Parikh +1
The proliferation of large language models has driven demand for long-context inference on resource-constrained edge platforms. However, deploying these models on Neural Processing…
cs.AI2024
Benchmarking Edge AI Platforms for High-Performance ML Inference
Rakshith Jayanth, Neelesh Gupta, Viktor Prasanna
Edge computing's growing prominence, due to its ability to reduce communication latency and enable real-time processing, is promoting the rise of high-performance, heterogeneous Sy…