2 papers
cs.SD2026
NPUsper: Eliminating Redundant Computation for Real-Time Whisper on Mobile NPUs
Sihyeon Lee, Hojeong Lee, Sungwon Woo +3
We present NPUsper, a live transcription system that makes Whisper efficient on mobile NPUs by eliminating redundant computation. To avoid the heavy padding used by prior streaming…
cs.AI2024
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
Taeho Kim, Yanming Wang, Vatshank Chaturvedi +4
Fine-tuning pre-trained large language models (LLMs) with limited hardware presents challenges due to GPU memory constraints. Various distributed fine-tuning methods have been prop…