2 papers
cs.CV2026
Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving
Christopher Lang, Alexander Braun, Abhinav Valada
Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, exist…
cs.RO2026
Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model
Harisankar Babu, Benjamin Coors, Christopher Lang +3
Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a represent…