paper

Context-Dependent Affordance Computation in Vision-Language Models

arXiv:2603.04419

Abstract

We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs). Our primary study uses Qwen3-VL-30B-A3B ( scene-context pairs from COCO-2017: 479 images under 7 agentic personas), with a cross-model replication on LLaVA-1.5-13B. We demonstrate substantial affordance drift: mean Jaccard similarity between context conditions is (95% CI across images; prime pairs; ), indicating that more than 90% of lexical scene description is context-dependent; the LLaVA replication reproduces the effect (mean , 84% context-dependent). Sentence-level cosine similarity confirms drift at the semantic level (mean , 58.5% context-dependent). Stochastic baseline experiments ( inference runs across 4 temperatures and 5 seeds) confirm this reflects genuine context effects rather than generation noise: within-prime variance is substantially lower than cross-prime variance across all conditions. Tucker decomposition with bootstrap stability analysis ( resamples) reveals stable orthogonal latent factors: a "Culinary Manifold" isolated to chef contexts and an "Access Axis" spanning child-mobility contrasts. The gap between lexical (90%) and semantic (58.5%) measures indicates that surface vocabulary changes more than underlying meaning under context shifts. These findings suggest a direction for robotics: dynamic, query-dependent ontological projection (JIT Ontology) rather than static world modeling. We do not claim to establish processing order or architectural primacy; such claims require internal representational analysis beyond output behavior.

33 pages, 13 tables, 3 figures. Code available at: https://github.com/studiofarzulla/semantic-vision

Context-Dependent Affordance Computation in Vision-Language Models · wovepaper