3 papers
cs.RO2026
ART-VS: Adaptive Resolution Tiling for Vision Transformer Visual Servoing
Alessandro Scherl, Bernhard Neuberger, Simon Schwaiger +3
Visual servoing with self-supervised Vision Transformer (ViT) features enables training-free robotic positioning with strong generalization, but faces a fundamental trade-off betwe…
cs.RO2026
ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks
Simon Schwaiger, David Seyser, Alessandro Scherl +2
Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings have shown to be noisy and lack spatial consistency. This is problem…
cs.RO2025
ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing
Alessandro Scherl, Stefan Thalhammer, Bernhard Neuberger +2
Visual servoing enables robots to precisely position their end-effector relative to a target object. While classical methods rely on hand-crafted features and thus are universally…