activity
20242026
collaborators

10 papers

cs.CV2026

When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection

Jihyeon Kim, Sohee Kim, Soosan Lee +3

Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularly in person-centric and partia…

cs.CV2026

AdaMerge: Salience-Aware Adaptive Token Merging for Training-Free Acceleration of Vision Transformers

Semi Lee, Hyejin Go, Hyesong Choi

The quadratic cost of self-attention in Vision Transformers (ViTs) constitutes a fundamental bottleneck for practical deployment, motivating a vibrant line of research on token red…

cs.CV2026

The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP

Kahyeon Nam, Hyesong Choi

Deploying Vision-Language Models on resource-constrained hardware typically requires INT8 quantization, but in joint-embedding architectures such as CLIP this introduces a failure…

cs.CV2026

What Does the Caption Really Say? Counterfactual Phrase Intervention for Compositional Data Selection in Vision-Language Pretraining

Hyejin Go, Semi Lee, Hyesong Choi

CLIP-style contrastive pretraining typically curates web-scale image-text pairs using sample-level filtering signals, often based on pair-level alignment. We show that this signal…

cs.CV2026

Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision

Jiyeong Kim, Yerim So, Hyesong Choi +2

Unified Multimodal Models (UMMs) have emerged as a promising paradigm that integrates multimodal understanding and generation within a unified modeling framework. However, current…

cs.CV2025

RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo

Jueun Ko, Hyewon Park, Hyesong Choi +1

Stereo Depth Estimation in real-world environments poses significant challenges due to dynamic domain shifts, sparse or unreliable supervision, and the high cost of acquiring dense…