8 papers
IMMSched: Interruptible Multi-DNN Scheduling via Parallel Multi-Particle Optimizing Subgraph Isomorphism
Boran Zhao, Hetian Liu, Zihang Yuan +4
The growing demand for multi-DNN workloads with unpredictable task arrival times has highlighted the need for interruptible scheduling on edge accelerators. However, existing preem…
FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
Haiduo Huang, Jiangcheng Song, Wenzhe Zhao +1
Speculative decoding accelerates generation by drafting candidates and verifying them in parallel, yet state-of-the-art drafters (e.g., EAGLE) still require N sequential passes to…
IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
Boran Zhao, Zihang Yuan, Yanbin Hu +5
Deploying deep neural network (DNN) accelerators with Layer Temporal Scheduling (LTS) often incurs significant overheads (e.g., energy and latency), as intermediate activations mus…
AdapSNE: Adaptive Fireworks-Optimized and Entropy-Guided Dataset Sampling for Edge DNN Training
Boran Zhao, Hetian Liu, Zihang Yuan +5
Training deep neural networks (DNNs) directly on edge devices has attracted increasing attention, as it offers promising solutions to challenges such as domain adaptation and priva…
SparseMap: A Sparse Tensor Accelerator Framework Based on Evolution Strategy
Boran Zhao, Haiming Zhai, Zihang Yuan +4
The growing demand for sparse tensor algebra (SpTA) in machine learning and big data has driven the development of various sparse tensor accelerators. However, most existing manual…
NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds
Boran Zhao, Haiduo Huang, Qiwei Dang +3
Training deep neural networks (DNNs) on edge devices has attracted increasing attention due to its potential to address challenges related to domain adaptation and privacy preserva…