activity
20242026
most citedSalience-Based Adaptive Masking: Revisiting Token Dynamics for Enhanced Pre-training

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision

Jiyeong Kim, Yerim So, Hyesong Choi +2

Unified Multimodal Models (UMMs) have emerged as a promising paradigm that integrates multimodal understanding and generation within a unified modeling framework. However, current…

cs.CV2025

RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo

Jueun Ko, Hyewon Park, Hyesong Choi +1

Stereo Depth Estimation in real-world environments poses significant challenges due to dynamic domain shifts, sparse or unreliable supervision, and the high cost of acquiring dense…

cs.CV2025

TADFormer : Task-Adaptive Dynamic Transformer for Efficient Multi-Task Learning

Seungmin Baek, Soyul Lee, Hayeon Jo +2

Transfer learning paradigm has driven substantial advancements in various vision tasks. However, as state-of-the-art models continue to grow, classical full fine-tuning often becom…

cs.CV2024

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models

Hyesong Choi, Daeun Kim, Sungmin Cha +2

In this work, we dive deep into the impact of additive noise in pre-training deep networks. While various methods have attempted to use additive noise inspired by the success of la…

cs.CV2024

MaDis-Stereo: Enhanced Stereo Matching via Distilled Masked Image Modeling

Jihye Ahn, Hyesong Choi, Soomin Kim +1

In stereo matching, CNNs have traditionally served as the predominant architectures. Although Transformer-based stereo models have been studied recently, their performance still la…

cs.CV2024

UniTT-Stereo: Unified Training of Transformer for Enhanced Stereo Matching

Soomin Kim, Hyesong Choi, Jihye Ahn +1

Unlike other vision tasks where Transformer-based approaches are becoming increasingly common, stereo depth estimation is still dominated by convolution-based approaches. This is m…