4 papers
Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers
Dong Hoon Lee, Seunghoon Hong
Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio. Vari…
Universal Few-Shot Spatial Control for Diffusion Models
Kiet T. Nguyen, Chanhyuk Lee, Donggyun Kim +2
Spatial conditioning in pretrained text-to-image diffusion models has significantly improved fine-grained control over the structure of generated images. However, existing control…
Disentangled Representation Learning via Modular Compositional Bias
Whie Jung, Dong Hoon Lee, Seunghoon Hong
Recent disentangled representation learning (DRL) methods heavily rely on factor specific strategies-either learning objectives for attributes or model architectures for objects-to…
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
Dong Hoon Lee, Seunghoon Hong
Recent token reduction methods for Vision Transformers (ViTs) incorporate token merging, which measures the similarities between token embeddings and combines the most similar pair…