collaborators

8 papers

cs.AR2026

IMMSched: Interruptible Multi-DNN Scheduling via Parallel Multi-Particle Optimizing Subgraph Isomorphism

Boran Zhao, Hetian Liu, Zihang Yuan +4

The growing demand for multi-DNN workloads with unpredictable task arrival times has highlighted the need for interruptible scheduling on edge accelerators. However, existing preem…

cs.LG2025

FastEagle: Cascaded Drafting for Accelerating Speculative Decoding

Haiduo Huang, Jiangcheng Song, Wenzhe Zhao +1

Speculative decoding accelerates generation by drafting candidates and verifying them in parallel, yet state-of-the-art drafters (e.g., EAGLE) still require N sequential passes to…

cs.DC2025

IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism

Boran Zhao, Zihang Yuan, Yanbin Hu +5

Deploying deep neural network (DNN) accelerators with Layer Temporal Scheduling (LTS) often incurs significant overheads (e.g., energy and latency), as intermediate activations mus…

cs.LG2025

AdapSNE: Adaptive Fireworks-Optimized and Entropy-Guided Dataset Sampling for Edge DNN Training

Boran Zhao, Hetian Liu, Zihang Yuan +5

Training deep neural networks (DNNs) directly on edge devices has attracted increasing attention, as it offers promising solutions to challenges such as domain adaptation and priva…

cs.LG2025

SparseMap: A Sparse Tensor Accelerator Framework Based on Evolution Strategy

Boran Zhao, Haiming Zhai, Zihang Yuan +4

The growing demand for sparse tensor algebra (SpTA) in machine learning and big data has driven the development of various sparse tensor accelerators. However, most existing manual…

cs.LG2025

NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds

Boran Zhao, Haiduo Huang, Qiwei Dang +3

Training deep neural networks (DNNs) on edge devices has attracted increasing attention due to its potential to address challenges related to domain adaptation and privacy preserva…