activity
20182026
most citedCharacter Region Awareness for Text Detection

58 citations · 175 across the 28 of their papers we have counts for

collaborators
Showing cs.CVShow all

35 papers · 1 filter

cs.CV2026

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Hyesong Choi, Daeun Kim, Song Park +5

Data augmentation is fundamental to training modern deep vision and multimodal models. While individual methods, such as RandAug, CutMix, Mixup, RandErase, and DropPath, offer stro…

cs.CV2026

Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Models

Hyundong Jin, Dongyoon Han, Eunwoo Kim

Continual unlearning poses the challenge of enabling large vision-language models to selectively refuse specific image-instruction pairs in response to sequential deletion requests…

cs.CV2026

Grounding World Simulation Models in a Real-World Metropolis

Junyoung Seo, Hyunwook Choi, Minkyung Kwon +10

What if a world simulation model could render not an imagined environment but a city that actually exists? Prior generative world models synthesize visually plausible yet artificia…

cs.CV2026

On the Reliability of Cue Conflict and Beyond

Pum Jun Kim, Seung-Ah Lee, Seongho Park +2

Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes. The cue-conflict benchmark has been influential in pro…

cs.CV2025

RL makes MLLMs see better than SFT

Junha Song, Sangdoo Yun, Dongyoon Han +2

A dominant assumption in Multimodal Language Model (MLLM) research is that its performance is largely inherited from the LLM backbone, given its immense parameter scale and remarka…

cs.CV2025

Exploring Conditions for Diffusion models in Robotic Control

Heeseong Shin, Byeongho Heo, Dongyoon Han +2

While pre-trained visual representations have significantly advanced imitation learning, they are often task-agnostic as they remain frozen during policy learning. In this work, we…