1 paper · 1 filter
Sangyun Chung, Youngjoon Yu, Se Yeon Kim +2
Large-scale Vision-Language Models (VLMs) have achieved notable progress in aligning visual inputs with text. However, their ability to deeply understand the unique physical proper…