2 papers
cs.LG2026
FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward
Sriman Achanta
Autoregressive decode repeatedly streams a growing KV cache, making attention a major cost at long context. Existing high-performance kernels use online softmax, which discovers a…
eess.IV2024
Brain Tumor Segmentation Based on Deep Learning, Attention Mechanisms, and Energy-Based Uncertainty Prediction
Zachary Schwehr, Sriman Achanta
Brain tumors are one of the deadliest forms of cancer with a mortality rate of over 80%. A quick and accurate diagnosis is crucial to increase the chance of survival. However, in m…