1 paper · 1 filter
Dawei Dai, Xu Long, Li Yutang +2
Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individua…