5 papers
Structured State-Space Regularization for Generation-Friendly Image Tokenization
Jinsung Lee, Jaemin Oh, Namhun Kim +3
Image tokenizers play a central role in modern generative models, where the structure of the latent space critically determines the downstream generation performance. A key but und…
Personalized Federated Learning for Gradient Alignment
Dongwon Kim, Gyuejeong Lee
Personalized federated learning (pFL) aims to adapt models to client specific data distributions, yet it often fails to reliably preserve personalized information. Local training i…
Raon-Speech Technical Report
Beomsoo Kim, Changho Choi, Dohyun Kim +23
We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and generation, and Raon-SpeechChat,…
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
Dongwon Kim, Ju He, Qihang Yu +4
Image tokenizers form the foundation of modern text-to-image generative models but are notoriously difficult to train. Furthermore, most existing text-to-image models rely on large…
1.58-bit FLUX
Chenglin Yang, Celong Liu, Xueqing Deng +4
We present 1.58-bit FLUX, the first successful approach to quantizing the state-of-the-art text-to-image generation model, FLUX.1-dev, using 1.58-bit weights (i.e., values in {-1,…