2 papers
cs.CV2026
Hierarchical Pre-Training of Vision Encoders with Large Language Model
Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai +2
The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks. However, existing approaches often tr…
cs.LG2026
Learning to Select Visual In-Context Demonstrations
Eugene Lee, Yu-Chi Lin, Jiajie Diao
Multimodal Large Language Models (MLLMs) adapt to visual tasks via in-context learning (ICL), which relies heavily on demonstration quality. The dominant demonstration selection st…