3 papers
cs.AR2026
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
Ma Zirui, Fan Zhihua, Li Wenxing +4
Speculative decoding enhances the inference efficiency of large language models (LLMs) by generating drafts using a small draft language model (DLM) and verifying them in batches w…
cs.AR2025
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
Shantian Qin, Ziqing Qiang, Zhihua Fan +4
Multimodal Transformers are emerging artificial intelligence (AI) models designed to process a mixture of signals from diverse modalities. Digital computing-in-memory (CIM) archite…
cs.AR2024
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
Haibin Wu, Wenming Li, Kai Yan +9
Recent neural networks (NNs) with self-attention exhibit competitiveness across different AI domains, but the essential attention mechanism brings massive computation and memory de…