activity
20192026
most citedYOLOv6: A Single-Stage Object Detection Framework for Industrial Applications

1.8k citations · 1.9k across the 12 of their papers we have counts for

collaborators
Showing 2022 · cs.CVShow all

6 papers · 2 filters

cs.CV2022

Uncertainty-Aware Image Captioning

Zhengcong Fei, Mingyuan Fan, Li Zhu +3

It is well believed that the higher uncertainty in a word of the caption, the more inter-correlated context information is required to determine it. However, current image captioni…

cs.CV2022★ 1 cited

Meta-Ensemble Parameter Learning

Zhengcong Fei, Shuman Tian, Junshi Huang +2

Ensemble of machine learning models yields improved performance as well as robustness. However, their memory requirements and inference costs can be prohibitively high. Knowledge d…

cs.CV2022★ 1.8k cited

YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications

Chuyi Li, Lulu Li, Hongliang Jiang +15

For years, the YOLO series has been the de facto industry-level standard for efficient object detection. The YOLO community has prospered overwhelmingly to enrich its use in a mult…

cs.CV2022

PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative Grounding

Zihan Ding, Zi-han Ding, Tianrui Hui +4

Panoptic Narrative Grounding (PNG) is an emerging task whose goal is to segment visual objects of things and stuff categories described by dense narrative captions of a still image…

cs.CV2022★ 1 cited

Efficient Modeling of Future Context for Image Captioning

Zhengcong Fei, Junshi Huang, Xiaoming Wei +1

Existing approaches to image captioning usually generate the sentence word-by-word from left to right, with the constraint of conditioned on local context including the given image…

cs.CV2022★ 3 cited

Language-Bridged Spatial-Temporal Interaction for Referring Video Object Segmentation

Zihan Ding, Tianrui Hui, Junshi Huang +3

Referring video object segmentation aims to predict foreground labels for objects referred by natural language expressions in videos. Previous methods either depend on 3D ConvNets…