95 citations · 114 across the 2 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 19 cited
Caption Anything: Interactive Image Description with Diverse Multimodal Controls
Teng Wang, Jinrui Zhang, Junjie Fei +5
Controllable image captioning is an emerging multimodal topic that aims to describe the image with natural language following human purpose, , looking at the specifi…
cs.CV2023★ 95 cited
Track Anything: Segment Anything Meets Videos
Jinyu Yang, Mingqi Gao, Zhe Li +3
Recently, the Segment Anything Model (SAM) gains lots of attention rapidly due to its impressive segmentation performance on images. Regarding its strong ability on image segmentat…
cs.CV2018
Multi-scale Location-aware Kernel Representation for Object Detection
Hao Wang, Qilong Wang, Mingqi Gao +2
Although Faster R-CNN and its variants have shown promising performance in object detection, they only exploit simple first-order representation of object proposals for final class…