3 papers
cs.CV2026
Teaching Vision-Language-Action Models What to See and Where to Look
Yuguang Yang, Canyu Chen, Zhewen Tan +10
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual q…
cs.CV2026
Noise-Robust Tiny Object Localization with Flows
Huixin Sun, Linlin Yang, Ronyu Chen +4
Despite significant advances in generic object detection, a persistent performance gap remains for tiny objects compared to normal-scale objects. We demonstrate that tiny objects a…
cs.CV2025
Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection
Yuguang Yang, Tongfei Chen, Haoyu Huang +5
Zero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Rece…