1 citations · 1 across the 7 of their papers we have counts for
6 papers · 1 filter
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
Hyojin Park, Yi Li, Janghoon Cho +8
Despite decades of work, surveillance still struggles in searching and reasoning about specific targets across long, multi-camera videos. Existing methods - tracking, retrieval, an…
Attention Guided Alignment in Efficient Vision-Language Models
Shweta Mahajan, Hoang Le, Hyojin Park +3
Large Vision-Language Models (VLMs) rely on effective multimodal alignment between pre-trained vision encoders and Large Language Models (LLMs) to integrate visual and textual info…
Generalized Contrastive Learning for Universal Multimodal Retrieval
Jungsoo Lee, Janghoon Cho, Hyojin Park +4
Despite their consistent performance improvements, cross-modal retrieval models (e.g., CLIP) show degraded performances with retrieving keys composed of fused image-text modality (…
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
Minho Park, Sunghyun Park, Jungsoo Lee +5
This paper addresses the challenge of data scarcity in semantic segmentation by generating datasets through text-to-image (T2I) generation models, reducing image acquisition and la…
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
Shubhankar Borse, Kartikeya Bhardwaj, Mohammad Reza Karimi Dastjerdi +8
Diffusion models are increasingly popular for generative tasks, including personalized composition of subjects and styles. While diffusion models can generate user-specified subjec…
PosSAM: Panoptic Open-vocabulary Segment Anything
Vibashan VS, Shubhankar Borse, Hyojin Park +4
In this paper, we introduce an open-vocabulary panoptic segmentation model that effectively unifies the strengths of the Segment Anything Model (SAM) with the vision-language CLIP…