2 papers
cs.LG2024
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
Shuzhang Zhong, Zebin Yang, Meng Li +3
Recent advancements in generative large language models (LLMs) have significantly boosted the performance in natural language processing tasks. However, their efficiency is hampere…
cs.LG2023
Memory-aware Scheduling for Complex Wired Networks with Iterative Graph Optimization
Shuzhang Zhong, Meng Li, Yun Liang +2
Memory-aware network scheduling is becoming increasingly important for deep neural network (DNN) inference on resource-constrained devices. However, due to the complex cell-level a…