25 citations · 34 across the 5 of their papers we have counts for
5 papers
Instance As Identity: A Generic Online Paradigm for Video Instance Segmentation
Feng Zhu, Zongxin Yang, Xin Yu +2
Modeling temporal information for both detection and tracking in a unified framework has been proved a promising solution to video instance segmentation (VIS). However, how to effe…
VMFormer: End-to-End Video Matting with Transformer
Jiachen Li, Vidit Goel, Marianna Ohanyan +3
Video matting aims to predict the alpha mattes for each frame from a given input video sequence. Recent solutions to video matting have been dominated by deep convolutional neural…
SiRi: A Simple Selective Retraining Mechanism for Transformer-based Visual Grounding
Mengxue Qu, Yu Wu, Wu Liu +5
In this paper, we investigate how to achieve better visual grounding with modern vision-language transformers, and propose a simple yet powerful Selective Retraining (SiRi) mechani…
Clicking Matters:Towards Interactive Human Parsing
Yutong Gao, Liqian Liang, Congyan Lang +3
In this work, we focus on Interactive Human Parsing (IHP), which aims to segment a human image into multiple human body parts with guidance from users' interactions. This new task…
Computational Baby Learning
Xiaodan Liang, Si Liu, Yunchao Wei +3
Intuitive observations show that a baby may inherently possess the capability of recognizing a new visual concept (e.g., chair, dog) by learning from only very few positive instanc…