8 papers
Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution
Bryan Sangwoo Kim, Jonghyun Park, Jong Chul Ye
Text-conditioned diffusion models have advanced image and video super-resolution by using prompts as semantic priors, and modern super-resolution pipelines typically rely on latent…
Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
Bryan Sangwoo Kim, Jeongsol Kim, Jong Chul Ye
Modern single-image super-resolution (SISR) models deliver photo-realistic results at the scale factors on which they are trained, but collapse when asked to magnify far beyond tha…
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
Ara Seo, Bryan Sangwoo Kim, Hyungjin Chung +1
Medical object detection suffers when a single detector is trained on mixed medical modalities (e.g., CXR, CT, MRI) due to heterogeneous statistics and disjoint representation spac…
FreeGuide: Training-Free Text-to-Video Alignment using Image LVLM
Jaemin Kim, Bryan Sangwoo Kim, Jong Chul Ye
Diffusion models have achieved impressive results in generative tasks for text-to-video (T2V) synthesis. However, achieving accurate text alignment in T2V generation remains challe…
Extreme Blind Image Restoration via Prompt-Conditioned Information Bottleneck
Hongeun Kim, Bryan Sangwoo Kim, Jong Chul Ye
Blind Image Restoration (BIR) methods have achieved remarkable success but falter when faced with Extreme Blind Image Restoration (EBIR), where inputs suffer from severe, compounde…
Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
Omkar Desai, Ziyang Jiao, Shuyi Pei +2
Input data preprocessing is a common bottleneck when concurrently training multimedia machine learning (ML) models in modern systems. To alleviate these bottlenecks and reduce the…