1 paper
Simon Schwaiger, David Seyser, Alessandro Scherl +2
Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings have shown to be noisy and lack spatial consistency. This is problem…