Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU
Victor J. B. Jung, Gagandeep Singh, Joseph Melber +3
The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-ch…
cs.DC2026
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
Enrico Russo, Mohamed Amine Hamdi, Alessandro Ottaviano +6
Deploying DNNs on System-on-Chips (SoC) with multiple heterogeneous acceleration engines is challenging, and the majority of deployment frameworks cannot fully exploit heterogeneit…