activity
20142023
most citedDeep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

652 citations · 1k across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20231 cited

Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints

Jiachen Li, Xinwei Shi, Feiyu Chen +9

Accurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersec…

cs.CV2021

Multi-modal 3D Human Pose Estimation with 2D Weak Supervision in Autonomous Driving

Jingxiao Zheng, Xinwei Shi, Alexander Gorban +9

3D human pose estimation (HPE) in autonomous vehicles (AV) differs from other use cases in many factors, including the 3D resolution and range of data, absence of dense depth maps,…

cs.LG201626 cited

Training and Evaluating Multimodal Word Embeddings with Large-scale Web Annotated Images

Junhua Mao, Jiajing Xu, Yushi Jing +1

In this paper, we focus on training and evaluating effective word embeddings with both text and visual information. More specifically, we introduce a large-scale dataset with 300 m…

cs.CV2014652 cited

Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

Junhua Mao, Wei Xu, Yi Yang +3

In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel image captions. It directly models the probability distribution of generating a w…

cs.CV2014370 cited

Explain Images with Multimodal Recurrent Neural Networks

Junhua Mao, Wei Xu, Yi Yang +2

In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel sentence descriptions to explain the content of images. It directly models the pr…