Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
Jingyi Wang, Jianzhong Ju, Jian Luan +1
Recent advances in large vision-language models (VLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches…
cs.CV2024
Context-Aware Aerial Object Detection: Leveraging Inter-Object and Background Relationships
Botao Ren, Botian Xu, Xue Yang +3
In most modern object detection pipelines, the detection proposals are processed independently given the feature map. Therefore, they overlook the underlying relationships between…
cs.CV2023
Feedback RoI Features Improve Aerial Object Detection
Botao Ren, Botian Xu, Tengyu Liu +2
Neuroscience studies have shown that the human visual system utilizes high-level feedback information to guide lower-level perception, enabling adaptation to signals of different c…