13 papers · 1 filter
MERIT: Multi-domain Efficient RAW Image Translation
Wenjun Huang, Shenghao Fu, Yian Jin +10
RAW images captured by different camera sensors exhibit substantial domain shifts due to varying spectral responses, noise characteristics, and tone behaviors, complicating their d…
Draft and Refine with Visual Experts
Sungheon Jeong, Ryozo Masukawa, Jihong Park +5
While recent Large Vision-Language Models (LVLMs) exhibit strong multimodal reasoning abilities, they often produce ungrounded or hallucinated responses because they rely too heavi…
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
Sanggeon Yun, Ryozo Masukawa, SungHeon Jeong +3
Vision-Language Models (VLMs) such as CLIP enable strong zero-shot recognition but suffer substantial degradation under distribution shifts. Test-Time Adaptation (TTA) aims to impr…
TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
Hanning Chen, Keyu Man, Kevin Zhu +8
Identifying and addressing performance anti-patterns in machine learning (ML) models is critical for efficient training and inference, but it typically demands deep expertise spann…
Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking
Wenjun Huang, Yang Ni, Hanning Chen +4
Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to localize an arbitrary number of targets based on a language expression and continuously track th…
Expanding Event Modality Applications through a Robust CLIP-Based Encoder
Sungheon Jeong, Hanning Chen, Sanggeon Yun +4
This paper introduces a powerful encoder that transfers CLIP`s capabilities to event-based data, enhancing its utility and expanding its applicability across diverse domains. While…