Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SelFusion: Self-distillation for Diffusion Language Models
Hyeongsoo Lim, Jinyoung Kim, Eunseo Seo +2
Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded generation quality limits practic…
cs.CL2024
LaDiMo: Layer-wise Distillation Inspired MoEfier
Sungyoon Kim, Youngjun Kim, Kihyo Moon +1
The advent of large language models has revolutionized natural language processing, but their increasing complexity has led to substantial training costs, resource demands, and env…