collaborators

18 papers

cs.LG2026

Activation Quantization of Vision Encoders Needs Prefixing Registers

Seunghyeon Kim, Taesun Yeom, Jinho Kim +3

Large pretrained vision encoders are central to multimodal intelligence, powering applications from on-device vision processing to vision-language models. Since these applications…

cs.CV2026

Learned Image Compression for Vision-Language-Action Models

Hyeonjun Kim, Jegwang Ryu, Sangbeom Ha +4

Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in b…

cs.CV2026

Understanding the Effects of Distractors on Reasoning Vision-Language Models

Jiyun Bae, Hyunjong Ok, Sangwoo Mo +1

How does irrelevant information (i.e., distractors) affect test-time scaling in vision-language models (VLMs)? Prior work on text-only language models has shown that textual distra…

cs.LG2026

Neural Weight Compression for Language Models

Jegwang Ryu, Minkyu Kim, Seungjun Shin +3

Efficient compression of language model weights is increasingly critical as model scale and deployment grow. Yet, most existing methods rely on handcrafted transforms and heuristic…

cs.LG2026

Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization

Taesun Yeom, Taehyeok Ha, Jaeho Lee

Feature learning strength (FLS), i.e., the inverse of the effective output scaling of a model, plays a critical role in shaping the optimization dynamics of neural nets. While its…

eess.IV2026

Multi-frame Restoration for 10 Hz Lissajous Confocal Laser Endomicroscopy

Minhee Lee, Sangyoon Lee, Jiwook Lee +4

Lissajous confocal laser endomicroscopy (CLE) is a promising solution for high-speed in vivo optical biopsy for handheld scenarios. However, Lissajous scanning traces a resonant tr…