4 papers
Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models
Hancheol Park, Geonho Lee, Tairen Piao +1
Mixture-of-Experts (MoE) models scale foundation models efficiently by activating only a subset of experts for each token, but their large number of expert parameters still makes q…
TabSODA: Tabular Diffusion based Imputation with Skip Pattern Detection and Ordinal Awareness
Yuyu Chen, Taehyo Kim, Hai Shu +1
Missing data imputation in large-scale surveys faces two challenges that are not well handled by current tabular diffusion methods. First, \emph{structural skips}, cells made inapp…
EdgeFusion: On-Device Text-to-Image Generation
Thibault Castells, Hyoung-Kyu Song, Tairen Piao +6
The intensive computational burden of Stable Diffusion (SD) for text-to-image generation poses a significant hurdle for its practical application. To tackle this challenge, recent…
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
Bo-Kyeong Kim, Geonmin Kim, Tae-Ho Kim +4
Structured pruning of modern large language models (LLMs) has emerged as a way of decreasing their high computational needs. Width pruning reduces the size of projection weight mat…