1 paper
Dawei Dai, Xu Long, Li Yutang +2
Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individua…