4 papers
SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
Yagiz Nalcakan, Hyeongjin Ju, Incheol Park +3
Vision foundation models (VFMs) pretrained on large-scale RGB data provide strong general-purpose representations, yet infrared perception, which is essential for robotics and driv…
RPT-SR: Regional Prior attention Transformer for infrared image Super-Resolution
Youngwan Jin, Incheol Park, Yagiz Nalcakan +3
General-purpose super-resolution models, particularly Vision Transformers, have achieved remarkable success but exhibit fundamental inefficiencies in common infrared imaging scenar…
Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
Junsu Kim, Naeun Kim, Jaeho Lee +3
The reasoning-based pose estimation (RPE) benchmark has emerged as a widely adopted evaluation standard for pose-aware multimodal large language models (MLLMs). Despite its signifi…
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
Youngwan Jin, Incheol Park, Hanbin Song +3
This paper proposes Pix2Next, a novel image-to-image translation framework designed to address the challenge of generating high-quality Near-Infrared (NIR) images from RGB inputs.…