activity
20242026
most citedPosSAM: Panoptic Open-vocabulary Segment Anything

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance

Hyojin Park, Yi Li, Janghoon Cho +8

Despite decades of work, surveillance still struggles in searching and reasoning about specific targets across long, multi-camera videos. Existing methods - tracking, retrieval, an…

cs.CV2025

Attention Guided Alignment in Efficient Vision-Language Models

Shweta Mahajan, Hoang Le, Hyojin Park +3

Large Vision-Language Models (VLMs) rely on effective multimodal alignment between pre-trained vision encoders and Large Language Models (LLMs) to integrate visual and textual info…

cs.CV2025

Generalized Contrastive Learning for Universal Multimodal Retrieval

Jungsoo Lee, Janghoon Cho, Hyojin Park +4

Despite their consistent performance improvements, cross-modal retrieval models (e.g., CLIP) show degraded performances with retrieving keys composed of fused image-text modality (…

cs.CV2025

CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation

Minho Park, Sunghyun Park, Jungsoo Lee +5

This paper addresses the challenge of data scarcity in semantic segmentation by generating datasets through text-to-image (T2I) generation models, reducing image acquisition and la…

cs.CV2025

SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization

Shubhankar Borse, Kartikeya Bhardwaj, Mohammad Reza Karimi Dastjerdi +8

Diffusion models are increasingly popular for generative tasks, including personalized composition of subjects and styles. While diffusion models can generate user-specified subjec…

cs.CV20241 cited

PosSAM: Panoptic Open-vocabulary Segment Anything

Vibashan VS, Shubhankar Borse, Hyojin Park +4

In this paper, we introduce an open-vocabulary panoptic segmentation model that effectively unifies the strengths of the Segment Anything Model (SAM) with the vision-language CLIP…