2 citations · 2 across the 3 of their papers we have counts for
4 papers
Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
Dohwan Ko, Ji Soo Lee, Minhyuk Choi +2
Text-Video Retrieval aims to find the most relevant text (or video) candidate given a video (or text) query from large-scale online databases. Recent work leverages multi-modal lar…
Efficient multi-view training for 3D Gaussian Splatting
Minhyuk Choi, Injae Kim, Hyunwoo J. Kim
3D Gaussian Splatting (3DGS) has emerged as a preferred choice alongside Neural Radiance Fields (NeRF) in inverse rendering due to its superior rendering speed. Currently, the comm…
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
Joonmyung Choi, Sanghyeok Lee, Jaewon Chu +2
Video Transformers have become the prevalent solution for various video downstream tasks with superior expressive power and flexibility. However, these video transformers suffer fr…
UP-NeRF: Unconstrained Pose-Prior-Free Neural Radiance Fields
Injae Kim, Minhyuk Choi, Hyunwoo J. Kim
Neural Radiance Field (NeRF) has enabled novel view synthesis with high fidelity given images and camera poses. Subsequent works even succeeded in eliminating the necessity of pose…