4 papers
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference
Pu Li, Jiawen Qi, Qinyu Chen
Deploying large language models (LLMs) on mobile devices increasingly relies on heterogeneous execution, yet no prior study has systematically characterized NPU effectiveness at th…
Two-Stage Adaptation for Non-Normative Speech Recognition: Revisiting Speaker-Independent Initialization for Personalization
Shan Jiang, Jiawen Qi, Chuanbing Huo +2
Personalizing automatic speech recognition (ASR) systems for non-normative speech, such as dysarthric and aphasic speech, is challenging. While speaker-specific fine-tuning (SS-FT)…
Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
Qinyu Chen, Jiawen Qi
Vision-Language Models (VLMs) deliver impressive performance in understanding visual content with language instructions. However, redundancy in vision tokens results in the degener…
DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference
Jiawen Qi, Chang Gao, Zhaochun Ren +1
Deploying Large Language Models (LLMs) on edge devices remains challenging due to their quadratically increasing computations with the sequence length. Existing studies for dynamic…