10 papers
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection
Jihyeon Kim, Sohee Kim, Soosan Lee +3
Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularly in person-centric and partia…
AdaMerge: Salience-Aware Adaptive Token Merging for Training-Free Acceleration of Vision Transformers
Semi Lee, Hyejin Go, Hyesong Choi
The quadratic cost of self-attention in Vision Transformers (ViTs) constitutes a fundamental bottleneck for practical deployment, motivating a vibrant line of research on token red…
The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP
Kahyeon Nam, Hyesong Choi
Deploying Vision-Language Models on resource-constrained hardware typically requires INT8 quantization, but in joint-embedding architectures such as CLIP this introduces a failure…
What Does the Caption Really Say? Counterfactual Phrase Intervention for Compositional Data Selection in Vision-Language Pretraining
Hyejin Go, Semi Lee, Hyesong Choi
CLIP-style contrastive pretraining typically curates web-scale image-text pairs using sample-level filtering signals, often based on pair-level alignment. We show that this signal…
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
Jiyeong Kim, Yerim So, Hyesong Choi +2
Unified Multimodal Models (UMMs) have emerged as a promising paradigm that integrates multimodal understanding and generation within a unified modeling framework. However, current…
RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo
Jueun Ko, Hyewon Park, Hyesong Choi +1
Stereo Depth Estimation in real-world environments poses significant challenges due to dynamic domain shifts, sparse or unreliable supervision, and the high cost of acquiring dense…