652 citations · 1.1k across the 7 of their papers we have counts for
7 papers · 1 filter
RefineVIS: Video Instance Segmentation with Temporal Attention Refinement
Andre Abrantes, Jiang Wang, Peng Chu +2
We introduce a novel framework called RefineVIS for Video Instance Segmentation (VIS) that achieves good object association between frames and accurate segmentation masks by iterat…
Adaptive Human Matting for Dynamic Videos
Chung-Ching Lin, Jiang Wang, Kun Luo +4
The most recent efforts in video matting have focused on eliminating trimap dependency since trimap annotations are expensive and trimap-based methods are less adaptable for real-t…
Binary Latent Diffusion
Ze Wang, Jiang Wang, Zicheng Liu +1
In this paper, we show that a binary latent space can be explored for compact yet expressive image representations. We model the bi-directional mappings between an image and the co…
Energy-Inspired Self-Supervised Pretraining for Vision Models
Ze Wang, Jiang Wang, Zicheng Liu +1
Motivated by the fact that forward and backward passes of a deep network naturally form symmetric mappings between input and output representations, we introduce a simple yet effec…
Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
Junhua Mao, Wei Xu, Yi Yang +3
In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel image captions. It directly models the probability distribution of generating a w…
Explain Images with Multimodal Recurrent Neural Networks
Junhua Mao, Wei Xu, Yi Yang +2
In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel sentence descriptions to explain the content of images. It directly models the pr…