2 papers
cs.LG2026
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
Maryam Hashemzadeh, Jerry Huang, Minseon Kim +2
The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmful prompts. While effective a…
cs.LG2025
Enhancing Variational Autoencoders with Smooth Robust Latent Encoding
Hyomin Lee, Minseon Kim, Sangwon Jang +2
Variational Autoencoders (VAEs) have played a key role in scaling up diffusion-based generative models, as in Stable Diffusion, yet questions regarding their robustness remain larg…