2 papers
cs.MM2026
When Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception
Guangyuan Dong, Chuang Liu, Haoyu Wang +10
Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retri…
cs.RO2026
Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics
Ran Chen, Jiaxing Ren, Zhikun Zhang +3
Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their…