7 papers · 1 filter
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
Gyuwon Han, Young Kyun Jang, Chanho Eom
Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing…
Visual Representation Alignment for Multimodal Large Language Models
Heeji Yoon, Jaewoo Jung, Junwan Kim +10
Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tas…
MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
Jeahun Sung, Changhyun Roh, Chanho Eom +1
Recent advances in portable imaging have made camera-based screen capture ubiquitous. Unfortunately, frequency aliasing between the camera's color filter array (CFA) and the displa…
R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision
Weeyoung Kwon, Jeahun Sung, Minkyu Jeon +2
Neural rendering methods such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have achieved significant progress in photorealistic 3D scene reconstruction and nov…
Domain Generalization for Person Re-identification: A Survey Towards Domain-Agnostic Person Matching
Hyeonseo Lee, Juhyun Park, Jihyong Oh +1
Person Re-identification (ReID) aims to retrieve images of the same individual captured across non-overlapping camera views, making it a critical component of intelligent surveilla…
AceVFI: A Comprehensive Survey of Advances in Video Frame Interpolation
Dahyeon Kye, Changhyun Roh, Sukhun Ko +2
Video Frame Interpolation (VFI) is a core low-level vision task that synthesizes intermediate frames between existing ones while ensuring spatial and temporal coherence. Over the p…