37 citations · 128 across the 33 of their papers we have counts for
12 papers · 1 filter
Video Text Tracking With a Spatio-Temporal Complementary Model
Yuzhe Gao, Xing Li, Jiajian Zhang +5
Text tracking is to track multiple texts in a video,and construct a trajectory for each text. Existing methodstackle this task by utilizing the tracking-by-detection frame-work, i.…
Anomaly Discovery in Semantic Segmentation via Distillation Comparison Networks
Huan Zhou, Shi Gong, Yu Zhou +3
This paper aims to address the problem of anomaly discovery in semantic segmentation. Our key observation is that semantic classification plays a critical role in existing approach…
A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-language Model
Mengde Xu, Zheng Zhang, Fangyun Wei +4
Recently, open-vocabulary image classification by vision language pre-training has demonstrated incredible achievements, that the model can classify arbitrary categories without se…
EPNet++: Cascade Bi-directional Fusion for Multi-Modal 3D Object Detection
Zhe Liu, Tengteng Huang, Bingling Li +3
Recently, fusing the LiDAR point cloud and camera image to improve the performance and robustness of 3D object detection has received more and more attention, as these two modaliti…
SeqFormer: Sequential Transformer for Video Instance Segmentation
Junfeng Wu, Yi Jiang, Song Bai +2
In this work, we present SeqFormer for video instance segmentation. SeqFormer follows the principle of vision transformer that models instance relationships among video frames. Nev…
SPTS: Single-Point Text Spotting
Dezhi Peng, Xinyu Wang, Yuliang Liu +9
Existing scene text spotting (i.e., end-to-end text detection and recognition) methods rely on costly bounding box annotations (e.g., text-line, word-level, or character-level boun…