Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices
Alejandro Ruiz y Mesa, Guilherme Korol, Moritz Riesterer +2
LLM deployment on resource-constrained edge devices faces severe latency constraints, particularly in real-time applications where delayed responses can compromise safety or usabil…
cs.LG2025
Leveraging Stochastic Depth Training for Adaptive Inference
Guilherme Korol, Antonio Carlos Schneider Beck, Jeronimo Castrillon
Dynamic DNN optimization techniques such as layer-skipping offer increased adaptability and efficiency gains but can lead to i) a larger memory footprint as in decision gates, ii)…