activity
20162022
most citedSemi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution Matching

67 citations · 193 across the 23 of their papers we have counts for

collaborators

26 papers

cs.CV20212 cited

Survey: Transformer based Video-Language Pre-training

Ludan Ruan, Qin Jin

Inspired by the success of transformer-based pre-training methods on natural language tasks and further computer vision tasks, researchers have begun to apply transformer to video…

cs.CV2021

Product-oriented Machine Translation with Cross-modal Cross-lingual Pre-training

Yuqing Song, Shizhe Chen, Qin Jin +3

Translating e-commercial product descriptions, a.k.a product-oriented machine translation (PMT), is essential to serve e-shoppers all over the world. However, due to the domain spe…

cs.CV20211 cited

Question-controlled Text-aware Image Captioning

Anwen Hu, Shizhe Chen, Qin Jin

For an image with multiple scene texts, different people may be interested in different text information. Current text-aware image captioning models are not able to generate distin…

cs.CV202120 cited

ICECAP: Information Concentrated Entity-aware Image Captioning

Anwen Hu, Shizhe Chen, Qin Jin

Most current image captioning systems focus on describing general image content, and lack background knowledge to deeply understand the image, such as exact named entities or concr…

cs.CL20215 cited

MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in Conversation

Jingwen Hu, Yuchen Liu, Jinming Zhao +1

Emotion recognition in conversation (ERC) is a crucial component in affective dialogue systems, which helps the system understand users' emotions and generate empathetic responses.…

cs.CV20211 cited

Team RUC_AIM3 Technical Report at ActivityNet 2021: Entities Object Localization

Ludan Ruan, Jieting Chen, Yuqing Song +2

Entities Object Localization (EOL) aims to evaluate how grounded or faithful a description is, which consists of caption generation and object grounding. Previous works tackle this…