3 citations · 3 across the 1 of their papers we have counts for
2 papers
cs.CV2025★ 3 cited
Maximizing the Position Embedding for Vision Transformers with Global Average Pooling
Wonjun Lee, Bumsub Ham, Suhyun Kim
In vision transformers, position embedding (PE) plays a crucial role in capturing the order of tokens. However, in vision transformer structures, there is a limitation in the expre…
cs.CV2025
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
Wonjun Lee, Doehyeon Lee, Eugene Choi +5
Current Vision Language Models (VLMs) remain vulnerable to malicious prompts that induce harmful outputs. Existing safety benchmarks for VLMs primarily rely on automated evaluation…