4 papers
Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs
Seungwoo Jung, Dohyeok Kwon, Seungmin Cha +4
Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to redu…
LSR-Net: A Lightweight and Strong Robustness Network for Bearing Fault Diagnosis in Noise Environment
Junseok Lee, Jihye Shin, Sangyong Lee +1
Rotating bearings play an important role in modern industries, but have a high probability of occurrence of defects because they operate at high speed, high load, and poor operatin…
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
Junseok Lee, Nahun Kim, Sangyong Lee +1
Knowledge distillation (KD) is one of the most effective paradigms for compressing large-scale foundation models into deployable architectures. In the context of Automatic Speech R…
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
Junseok Lee, Chang-Jae Chun
Scaling Multimodal Large Language Models (MLLMs) to long-form speech is bottlenecked by the explosive growth of input tokens. Existing speech-language models project high-frame-rat…