activity
20242026
collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

MERIT: Multi-domain Efficient RAW Image Translation

Wenjun Huang, Shenghao Fu, Yian Jin +10

RAW images captured by different camera sensors exhibit substantial domain shifts due to varying spectral responses, noise characteristics, and tone behaviors, complicating their d…

cs.CV2026

Draft and Refine with Visual Experts

Sungheon Jeong, Ryozo Masukawa, Jihong Park +5

While recent Large Vision-Language Models (LVLMs) exhibit strong multimodal reasoning abilities, they often produce ungrounded or hallucinated responses because they rely too heavi…

cs.CV2026

Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models

Sanggeon Yun, Ryozo Masukawa, SungHeon Jeong +3

Vision-Language Models (VLMs) such as CLIP enable strong zero-shot recognition but suffer substantial degradation under distribution shifts. Test-Time Adaptation (TTA) aims to impr…

cs.CV2025

TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models

Hanning Chen, Keyu Man, Kevin Zhu +8

Identifying and addressing performance anti-patterns in machine learning (ML) models is critical for efficient training and inference, but it typically demands deep expertise spann…

cs.CV2025

Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking

Wenjun Huang, Yang Ni, Hanning Chen +4

Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to localize an arbitrary number of targets based on a language expression and continuously track th…

cs.CV2025

Expanding Event Modality Applications through a Robust CLIP-Based Encoder

Sungheon Jeong, Hanning Chen, Sanggeon Yun +4

This paper introduces a powerful encoder that transfers CLIP`s capabilities to event-based data, enhancing its utility and expanding its applicability across diverse domains. While…