2 citations · 5 across the 17 of their papers we have counts for
10 papers · 1 filter
MERIT: Multi-domain Efficient RAW Image Translation
Wenjun Huang, Shenghao Fu, Yian Jin +10
RAW images captured by different camera sensors exhibit substantial domain shifts due to varying spectral responses, noise characteristics, and tone behaviors, complicating their d…
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
Sanggeon Yun, Ryozo Masukawa, SungHeon Jeong +3
Vision-Language Models (VLMs) such as CLIP enable strong zero-shot recognition but suffer substantial degradation under distribution shifts. Test-Time Adaptation (TTA) aims to impr…
Draft and Refine with Visual Experts
Sungheon Jeong, Ryozo Masukawa, Jihong Park +5
While recent Large Vision-Language Models (LVLMs) exhibit strong multimodal reasoning abilities, they often produce ungrounded or hallucinated responses because they rely too heavi…
LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation
Hanning Chen, Yang Ni, Wenjun Huang +4
Large Vision Language Models (LVLMs) have been widely adopted to guide vision foundation models in performing reasoning segmentation tasks, achieving impressive performance. Howeve…
Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking
Wenjun Huang, Yang Ni, Hanning Chen +4
Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to localize an arbitrary number of targets based on a language expression and continuously track th…
Expanding Event Modality Applications through a Robust CLIP-Based Encoder
Sungheon Jeong, Hanning Chen, Sanggeon Yun +4
This paper introduces a powerful encoder that transfers CLIP`s capabilities to event-based data, enhancing its utility and expanding its applicability across diverse domains. While…