6 papers
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
Kanghyun Baek, Jaihyun Lew, Chaehun Shin +2
Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omission, where specified objects…
HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
Yongsung Kim, Wooseok Song, Jaihyun Lew +3
Visual Geometry Grounded Transformer (VGGT) has advanced 3D vision, yet its global attention layers suffer from quadratic computational costs that hinder scalability. Several spars…
Causality-Aware Contrastive Learning for Robust Multivariate Time-Series Anomaly Detection
HyunGi Kim, Jisoo Mok, Dongjun Lee +3
Utilizing the complex inter-variable causal relationships within multivariate time-series provides a promising avenue toward more robust and reliable multivariate time-series anoma…
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
Jaihyun Lew, Soohyuk Jang, Jaehoon Lee +6
Transformers, a groundbreaking architecture proposed for Natural Language Processing (NLP), have also achieved remarkable success in Computer Vision. A cornerstone of their success…
Disentangled Motion Modeling for Video Frame Interpolation
Jaihyun Lew, Jooyoung Choi, Chaehun Shin +2
Video Frame Interpolation (VFI) aims to synthesize intermediate frames between existing frames to enhance visual smoothness and quality. Beyond the conventional methods based on th…
Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
Sanghyeob Song, Jaihyun Lew, Hyemi Jang +1
Estimating the homography between two images is crucial for mid- or high-level vision tasks, such as image stitching and fusion. However, using supervised learning methods is often…