3 papers
cs.CV2024
UniTT-Stereo: Unified Training of Transformer for Enhanced Stereo Matching
Soomin Kim, Hyesong Choi, Jihye Ahn +1
Unlike other vision tasks where Transformer-based approaches are becoming increasingly common, stereo depth estimation is still dominated by convolution-based approaches. This is m…
cs.CV2024
SG-MIM: Structured Knowledge Guided Efficient Pre-training for Dense Prediction
Sumin Son, Hyesong Choi, Dongbo Min
Masked Image Modeling (MIM) techniques have redefined the landscape of computer vision, enabling pre-trained models to achieve exceptional performance across a broad spectrum of ta…
cs.CV2022
Sequential Cross Attention Based Multi-task Learning
Sunkyung Kim, Hyesong Choi, Dongbo Min
In multi-task learning (MTL) for visual scene understanding, it is crucial to transfer useful information between multiple tasks with minimal interferences. In this paper, we propo…