robotics

Semantic Anchoring for Robotic Action Representations

arXiv:2607.13597

summary

The paper studies how fine‑tuning vision‑language‑action models for robots can degrade the semantic structure of their action representations, and proposes a plug‑and‑play anchoring technique that preserves this structure to improve task performance and generalization.

Abstract

Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. A fundamental question therefore arises: what constitutes a good action representation? Inspired by the mirror neuron theory's insight that observation and execution share an intention-level encoding, we examine whether a robot's action representations preserve the semantic structure captured by pretrained encoders. Systematic probing confirms that this structure erodes during finetuning, and that its quality synchronizes with both task success and out-of-distribution generalization. We further introduce a plug-and-play method that anchors action representations to a semantic manifold while decomposing representations into a shared semantic channel and a private channel, all discarded at inference, leaving the deployed model unchanged. Validated on different VLA backbones across simulation and real-world benchmarks, our method yields up to +18.7% on real-world in-distribution tasks and +21.5% on out-of-distribution generalization.

Project Page: https://xy02-05.github.io/SemanticMN

Topics & keywords

#action representation#vision-language models#semantic anchoring#robot learning#out-of-distribution generalizationvision-language-actionsemantic manifoldshared semantic channelplug-and-play anchoringfine-tuning degradation
Semantic Anchoring for Robotic Action Representations · wovepaper