1 paper · 1 filter
Animesh Tripathy, Aswanth Krishnan
Letting a vision-language model (VLM) think longer at test time has driven much recent progress. A natural way to bring this to spatial grounding is visual self-correction: the mod…