6 papers
Training Text-to-Molecule Models with Context-Aware Tokenization
Seojin Kim, Hyeontae Song, Jaehyun Nam +1
Recently, text-to-molecule models have shown great potential across various chemical applications, e.g., drug-discovery. These models adapt language models to molecular data by rep…
MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation
Sihyun Yu, Meera Hahn, Dan Kondratyuk +6
Diffusion models are successful for synthesizing high-quality videos but are limited to generating short clips (e.g., 2-10 seconds). Synthesizing sustained footage (e.g. over minut…
FontAdapter: Instant Font Adaptation in Visual Text Generation
Myungkyu Koo, Subin Kim, Sangkyung Kwak +3
Text-to-image diffusion models have significantly improved the seamless integration of visual text into diverse image contexts. Recent approaches further improve control over font…
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
Huiwon Jang, Sihyun Yu, Jinwoo Shin +2
Efficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop a tokenizer that can encode long…
Controllable Human Image Generation with Personalized Multi-Garments
Yisol Choi, Sangkyung Kwak, Sihyun Yu +2
We present BootComp, a novel framework based on text-to-image diffusion models for controllable human image generation with multiple reference garments. Here, the main bottleneck i…
Confidence-aware Denoised Fine-tuning of Off-the-shelf Models for Certified Robustness
Suhyeok Jang, Seojin Kim, Jinwoo Shin +1
The remarkable advances in deep learning have led to the emergence of many off-the-shelf classifiers, e.g., large pre-trained models. However, since they are typically trained on c…