3 citations · 3 across the 2 of their papers we have counts for
4 papers · 1 filter
VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Jiacheng Ruan, Wenzhen Yuan, Xian Gao +6
Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Re…
Salient Object Ranking with Position-Preserved Attention
Hao Fang, Daoxin Zhang, Yi Zhang +5
Instance segmentation can detect where the objects are in an image, but hard to understand the relationship between them. We pay attention to a typical relationship, relative salie…
Horizontal-to-Vertical Video Conversion
Tun Zhu, Daoxin Zhang, Yao Hu +4
Alongside the prevalence of mobile videos, the general public leans towards consuming vertical videos on hand-held devices. To revitalize the exposure of horizontal contents, we he…
Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
Xiongwei Wu, Daoxin Zhang, Jianke Zhu +1
Recent years have witnessed many exciting achievements for object detection using deep learning techniques. Despite achieving significant progresses, most existing detectors are de…