1 paper
Zesheng Yang, Xi Jiang, Bingzhang Hu +4
Current vision-language detection and grounding models predominantly focus on prompts with positive semantics and often struggle to accurately interpret and ground complex expressi…