3 papers
eess.IV2026
Cost-Aware Vision--Language Model Arbitration for Fabric Structure Recognition A Deployable Multi-Agent System
Chenwei Wang, Haochen Li, Shuk Ching Tang +3
Recognizing a fabric's structure is a prerequisite for translating textile-specific material information into structured digital form for downstream supply-chain systems. Pure CNN…
cs.RO2026
Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation
Zhiruo Zhou, Zelin Li, Xiwen Chen +4
Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enter…
cs.RO2026
Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation
Xiwen Chen, Zelin Li, Zhiruo Zhou +3
Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal r…