3 papers
cs.RO2026
Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models
Cong Xu, Ravi Sankar
End-to-end vision-language models (VLMs) bind visual competence to the scale of their language model: as the language model shrinks, perception and reasoning degrade together. We s…
cs.RO2026
Gating Before Commitment: Anticipating Intent Divergence to Prevent Post-Interaction Decision Failures in Autonomous Driving
Cong Xu, Ravi Sankar
Intent misinterpretation during vehicle interactions causes recurring planning failures. We study a decision layer in which a language-guided intent module reads structured descrip…
cs.RO2026
Evidence of Absence: Cross-Modal Abductive Risk Perception to Sustain World Models When Vision Fails
Cong Xu, Ravi Sankar
A structured world-state (entities, relations, context, and predictive cues) is designed to preserve prediction-critical content when perception degrades, but it presumes observati…