17 citations · 22 across the 4 of their papers we have counts for
4 papers · 1 filter
Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension
Huibin Zhang, Zhengkun Zhang, Yao Zhang +5
Procedural Multimodal Documents (PMDs) organize textual instructions and corresponding images step by step. Comprehending PMDs and inducing their representations for the downstream…
A General Framework for Learning Prosodic-Enhanced Representation of Rap Lyrics
Hongru Liang, Haozheng Wang, Qian Li +5
Learning and analyzing rap lyrics is a significant basis for many web applications, such as music recommendation, automatic music categorization, and music information retrieval, d…
Multi-modal Summarization for Video-containing Documents
Xiyan Fu, Jun Wang, Zhenglu Yang
Summarization of multimedia data becomes increasingly significant as it is the basis for many real-world applications, such as question answering, Web search, and so forth. Most ex…
JTAV: Jointly Learning Social Media Content Representation by Fusing Textual, Acoustic, and Visual Features
Hongru Liang, Haozheng Wang, Jun Wang +4
Learning social media content is the basis of many real-world applications, including information retrieval and recommendation systems, among others. In contrast with previous work…