46 citations · 99 across the 10 of their papers we have counts for
13 papers · 1 filter
SteerVTE: Seamless Video Text Editing with Style and Glyph Control
Kai Zeng, Moran Li, Zhengwei Wang +6
Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain,…
TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding
Zhengwei Wang, Qi She, Aljosa Smolic
Most of existing video action recognition models ingest raw RGB frames. However, the raw video stream requires enormous storage and contains significant temporal redundancy. Video…
3rd Place Solution to Google Landmark Recognition Competition 2021
Cheng Xu, Weimin Wang, Shuai Liu +6
In this paper, we show our solution to the Google Landmark Recognition 2021 Competition. Firstly, embeddings of images are extracted via various architectures (i.e. CNN-, Transform…
MT-ORL: Multi-Task Occlusion Relationship Learning
Panhe Feng, Qi She, Lei Zhu +7
Retrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of a…
Inter-intra Variant Dual Representations forSelf-supervised Video Recognition
Lin Zhang, Qi She, Zhengyang Shen +1
Contrastive learning applied to self-supervised representation learning has seen a resurgence in deep models. In this paper, we find that existing contrastive learning based soluti…
Learning the Superpixel in a Non-iterative and Lifelong Manner
Lei Zhu, Qi She, Bin Zhang +4
Superpixel is generated by automatically clustering pixels in an image into hundreds of compact partitions, which is widely used to perceive the object contours for its excellent c…