1 citations · 1 across the 6 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning
Xiaowen Sun, Matthias Kerzel, Mengdi Li +3
Vision-language models have demonstrated strong performance across robotic perception and instruction-following tasks. However, they still struggle with precise spatial reasoning,…
cs.CV2021
Visual Distant Supervision for Scene Graph Generation
Yuan Yao, Ao Zhang, Xu Han +5
Scene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer visi…