1 paper
Dawar Jyoti Deka, Amit Sethi, Syed Mohammad Ali
Vision-language models enable open-vocabulary object grounding through natural language queries, under the implicit assumption that semantically equivalent descriptions yield consi…