activity
20182026
most citedAdaptive Inference through Early-Exit Networks: Design, Challenges and Directions

106 citations · 127 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

WhiFlash: Accelerating Speculative Decoding with Token-Level Cross-Paradigm Routing

Young D. Kwon, Miles Williams, Rui Li +2

The autoregressive nature of large language models (LLMs) remains a significant bottleneck for inference, particularly in complex agentic workloads. While speculative decoding (SD)…

cs.LG2024★ 1 cited

Progressive Mixed-Precision Decoding for Efficient LLM Inference

Hao Mark Chen, Fuwen Tan, Alexandros Kouris +3

In spite of the great potential of large language models (LLMs) across various tasks, their deployment on resource-constrained devices remains challenging due to their excessive co…

cs.LG2022★ 2 cited

The Future of Consumer Edge-AI Computing

Stefanos Laskaridis, Stylianos I. Venieris, Alexandros Kouris +2

In the last decade, Deep Learning has rapidly infiltrated the consumer end, mainly thanks to hardware acceleration across devices. However, as we look towards the future, it is evi…

cs.LG2022★ 6 cited

Fluid Batching: Exit-Aware Preemptive Serving of Early-Exit Neural Networks on Edge NPUs

Alexandros Kouris, Stylianos I. Venieris, Stefanos Laskaridis +1

With deep neural networks (DNNs) emerging as the backbone in a multitude of computer vision tasks, their adoption in real-world applications broadens continuously. Given the abunda…

cs.LG2021★ 106 cited

Adaptive Inference through Early-Exit Networks: Design, Challenges and Directions

Stefanos Laskaridis, Alexandros Kouris, Nicholas D. Lane

DNNs are becoming less and less over-parametrised due to recent advances in efficient model design, through careful hand-crafted or NAS-based methods. Relying on the fact that not…