1 paper
Zhangxuan Gu, Zhuoer Xu, Haoxing Chen +3
Recent object detection approaches rely on pretrained vision-language models for image-text alignment. However, they fail to detect the Mobile User Interface (MUI) element since it…