2 papers
cs.LG2025
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra +1
Attention is the dominant source of latency during long-context LLM inference, an increasingly popular workload with reasoning models and RAG. We propose Kascade, a training-free s…
cs.LG2023
Entropy Aware Training for Fast and Accurate Distributed GNN
Dhruv Deshmukh, Gagan Raj Gupta, Manisha Chawla +2
Several distributed frameworks have been developed to scale Graph Neural Networks (GNNs) on billion-size graphs. On several benchmarks, we observe that the graph partitions generat…