From the 1 of 10 linked papers with an AI index.
10 papers
FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
Minguk Kang, Suha Kwak
FlashDecoder is a pure‑Transformer video decoder that converts latent representations to pixel frames in real time, using a rolling key‑value cache to keep computation and memory c…
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
Youngkil Song, Yoonjae Baek, Dongwon Kim +3
Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, coupling high-level reasoning with temporal groun…
ACID: Action Consistency via Inverse Dynamics for Planning with World Models
Gawon Seo, Dongwon Kim, Suha Kwak
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how…
Structured State-Space Regularization for Generation-Friendly Image Tokenization
Jinsung Lee, Jaemin Oh, Namhun Kim +3
Image tokenizers play a central role in modern generative models, where the structure of the latent space critically determines the downstream generation performance. A key but und…
TextME: Bridging Unseen Modalities Through Text Descriptions
Soyeon Hong, Jinchan Kim, Jaegook You +3
Expanding multimodal representations to novel modalities is constrained by reliance on large-scale paired datasets (e.g., text-image, text-audio, text-3D, text-molecule), which are…
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
Dongwon Kim, Ju He, Qihang Yu +4
Image tokenizers form the foundation of modern text-to-image generative models but are notoriously difficult to train. Furthermore, most existing text-to-image models rely on large…