1 paper · 1 filter
Guikun Chen, Jin Li, Wenguan Wang
Current approaches for open-vocabulary scene graph generation (OVSGG) use vision-language models such as CLIP and follow a standard zero-shot pipeline -- computing similarity betwe…