#model pruning
6 papers match
Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
Huiyuan Tian, Bonan Xu, Shijian Li
The paper investigates how sparse mixture-of-experts language models route tokens to multiple experts, showing that expert subspaces overlap substantially yet routing still selects…
VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment
Stephen Bauer, Sheila Seidel, Shanza Iftikhar +2
The paper introduces kiloVAD, an ultra‑tiny, CNN‑only voice activity detection model designed for edge devices, using standard Mel features, structured pruning with self‑distillati…
Post-Training Pruning for Diffusion Transformers
Chengzhi Hu, Xuewen Liu, Jing Zhang +3
The paper introduces DiT-Pruning, a post‑training pruning method tailored for Diffusion Transformers that uses a new energy‑based saliency metric and clustering‑aware granularity t…
Efficient Text-to-Audio Generation via Pruning
Arshdeep Singh, Yi Yuan, Yun Chen +2
The paper applies filter‑based pruning to the U‑Net backbone of the AudioLDM text‑to‑audio diffusion model, reducing most of its parameters and compute while preserving generation…
UMoE:Unlocking Every Expert in Domain-Specific Training
Xuefeng Li, Pengfei Liu
The paper introduces UMoE, a method that prunes low‑saliency experts and regrows new ones to better align a mixture‑of‑experts language model with a target domain before fine‑tunin…
An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals
Raktim Bhattacharya
The paper introduces an exact measurement tool that quantifies how selective state‑space models like Mamba use their internal modes, enabling precise prediction of pruning errors a…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.