most citedDeep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

652 citations · 1.1k across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV20231 cited

RefineVIS: Video Instance Segmentation with Temporal Attention Refinement

Andre Abrantes, Jiang Wang, Peng Chu +2

We introduce a novel framework called RefineVIS for Video Instance Segmentation (VIS) that achieves good object association between frames and accurate segmentation masks by iterat…

cs.CV2023

Adaptive Human Matting for Dynamic Videos

Chung-Ching Lin, Jiang Wang, Kun Luo +4

The most recent efforts in video matting have focused on eliminating trimap dependency since trimap annotations are expensive and trimap-based methods are less adaptable for real-t…

cs.CV2023

Binary Latent Diffusion

Ze Wang, Jiang Wang, Zicheng Liu +1

In this paper, we show that a binary latent space can be explored for compact yet expressive image representations. We model the bi-directional mappings between an image and the co…

cs.CV20231 cited

Energy-Inspired Self-Supervised Pretraining for Vision Models

Ze Wang, Jiang Wang, Zicheng Liu +1

Motivated by the fact that forward and backward passes of a deep network naturally form symmetric mappings between input and output representations, we introduce a simple yet effec…

cs.CV2014652 cited

Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

Junhua Mao, Wei Xu, Yi Yang +3

In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel image captions. It directly models the probability distribution of generating a w…

cs.CV2014370 cited

Explain Images with Multimodal Recurrent Neural Networks

Junhua Mao, Wei Xu, Yi Yang +2

In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel sentence descriptions to explain the content of images. It directly models the pr…