1 citations · 1 across the 8 of their papers we have counts for
11 papers · 1 filter
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
Jiyeong Kim, Yerim So, Hyesong Choi +2
Unified Multimodal Models (UMMs) have emerged as a promising paradigm that integrates multimodal understanding and generation within a unified modeling framework. However, current…
RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo
Jueun Ko, Hyewon Park, Hyesong Choi +1
Stereo Depth Estimation in real-world environments poses significant challenges due to dynamic domain shifts, sparse or unreliable supervision, and the high cost of acquiring dense…
TADFormer : Task-Adaptive Dynamic Transformer for Efficient Multi-Task Learning
Seungmin Baek, Soyul Lee, Hayeon Jo +2
Transfer learning paradigm has driven substantial advancements in various vision tasks. However, as state-of-the-art models continue to grow, classical full fine-tuning often becom…
Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models
Hyesong Choi, Daeun Kim, Sungmin Cha +2
In this work, we dive deep into the impact of additive noise in pre-training deep networks. While various methods have attempted to use additive noise inspired by the success of la…
MaDis-Stereo: Enhanced Stereo Matching via Distilled Masked Image Modeling
Jihye Ahn, Hyesong Choi, Soomin Kim +1
In stereo matching, CNNs have traditionally served as the predominant architectures. Although Transformer-based stereo models have been studied recently, their performance still la…
UniTT-Stereo: Unified Training of Transformer for Enhanced Stereo Matching
Soomin Kim, Hyesong Choi, Jihye Ahn +1
Unlike other vision tasks where Transformer-based approaches are becoming increasingly common, stereo depth estimation is still dominated by convolution-based approaches. This is m…