93 citations · 93 across the 1 of their papers we have counts for
2 papers
cs.MM2020★ 93 cited
Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation Learning
Ying Cheng, Ruize Wang, Zhihao Pan +2
When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlyi…
cs.CL2019
Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent Communication
Ruize Wang, Zhongyu Wei, Ying Cheng +5
Visual storytelling aims to generate a narrative paragraph from a sequence of images automatically. Existing approaches construct text description independently for each image and…