16 citations · 20 across the 8 of their papers we have counts for
6 papers · 1 filter
Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
Young Kyun Jang, Dat Huynh, Ashish Shah +2
Composed Image Retrieval (CIR) is a complex task that retrieves images using a query, which is configured with an image and a caption that describes desired modifications to that i…
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
Bo He, Hengduo Li, Young Kyun Jang +5
With the success of large language models (LLMs), integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However,…
Revisiting Kernel Temporal Segmentation as an Adaptive Tokenizer for Long-form Video Understanding
Mohamed Afham, Satya Narayan Shukla, Omid Poursaeed +3
While most modern video understanding models operate on short-range clips, real-world videos are often several minutes long with semantically consistent segments of variable length…
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning
Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed +4
We introduce Patch Aligned Contrastive Learning (PACL), a modified compatibility function for CLIP's contrastive loss, intending to train an alignment between the patch tokens of t…
Raising the Bar on the Evaluation of Out-of-Distribution Detection
Jishnu Mukhoti, Tsung-Yu Lin, Bor-Chun Chen +4
In image classification, a lot of development has happened in detecting out-of-distribution (OoD) data. However, most OoD detection methods are evaluated on a standard set of datas…
MixNorm: Test-Time Adaptation Through Online Normalization Estimation
Xuefeng Hu, Gokhan Uzunbas, Sirius Chen +4
We present a simple and effective way to estimate the batch-norm statistics during test time, to fast adapt a source model to target test samples. Known as Test-Time Adaptation, mo…