From the 1 of 12 linked papers with an AI index.
12 papers
Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval
Boseung Jeong, Taegyu Park, Donghyeon Kwon +2
At the heart of composed visual data retrieval is the fusion of a reference visual input and a textual modification into a single query. While current state-of-the-art methods util…
FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
Minguk Kang, Suha Kwak
FlashDecoder is a pure‑Transformer video decoder that converts latent representations to pixel frames in real time, using a rolling key‑value cache to keep computation and memory c…
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
Youngkil Song, Yoonjae Baek, Dongwon Kim +3
Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, coupling high-level reasoning with temporal groun…
ACID: Action Consistency via Inverse Dynamics for Planning with World Models
Gawon Seo, Dongwon Kim, Suha Kwak
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how…
Structured State-Space Regularization for Generation-Friendly Image Tokenization
Jinsung Lee, Jaemin Oh, Namhun Kim +3
Image tokenizers play a central role in modern generative models, where the structure of the latent space critically determines the downstream generation performance. A key but und…
TextME: Bridging Unseen Modalities Through Text Descriptions
Soyeon Hong, Jinchan Kim, Jaegook You +3
Expanding multimodal representations to novel modalities is constrained by reliance on large-scale paired datasets (e.g., text-image, text-audio, text-3D, text-molecule), which are…