2 papers
cs.CV2026
From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability
Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou
How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth pe…
cs.CV2026
Pretraining Objective Matters in Extreme Low-Data FGVC: A Backbone-Controlled Study
Alexander Hackett, Srikanth Thudumu, Ginny Fisher +1
Extreme low-data fine-grained classification is common in expert domains where labeling is expensive, yet practitioners still need principled guidance for selecting pretrained enco…