6 papers
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
Gyuwon Han, Young Kyun Jang, Chanho Eom
Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing…
Visual Representation Alignment for Multimodal Large Language Models
Heeji Yoon, Jaewoo Jung, Junwan Kim +10
Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tas…
MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
Jeahun Sung, Changhyun Roh, Chanho Eom +1
Recent advances in portable imaging have made camera-based screen capture ubiquitous. Unfortunately, frequency aliasing between the camera's color filter array (CFA) and the displa…
R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision
Weeyoung Kwon, Jeahun Sung, Minkyu Jeon +2
Neural rendering methods such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have achieved significant progress in photorealistic 3D scene reconstruction and nov…
Domain Generalization for Person Re-identification: A Survey Towards Domain-Agnostic Person Matching
Hyeonseo Lee, Juhyun Park, Jihyong Oh +1
Person Re-identification (ReID) aims to retrieve images of the same individual captured across non-overlapping camera views, making it a critical component of intelligent surveilla…
GOAL: Global-local Object Alignment Learning
Hyungyu Choi, Young Kyun Jang, Chanho Eom
Vision-language models like CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions because of thei…