1 paper · 1 filter
Aya Nakayama, Brian Wong, Yuji Nishimura +1
The "style trap" poses a significant challenge for Large Vision-Language Models (LVLMs), hindering robust semantic understanding across diverse visual styles, especially in in-cont…