2 papers
cs.LG2026
Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models
Bishmoy Paul, Youngmin Yi, Hoeseok Yang
Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perc…
cs.PF2024
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
Jiho Shin, Hoeseok Yang, Youngmin Yi
Leveraging sparsity is crucial for optimizing large language model inference. however, modern LLMs employing SiLU as their activation function exhibit minimal activation sparsity.…