Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Point What You Mean: Visually Grounded Instruction Policy
Hang Yu, Juntu Zhao, Yufeng Liu +9
Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especial…
cs.CV2023
Bi-directional Adapter for Multi-modal Tracking
Bing Cao, Junliang Guo, Pengfei Zhu +1
Due to the rapid development of computer vision, single-modal (RGB) object tracking has made significant progress in recent years. Considering the limitation of single imaging sens…