12 citations · 15 across the 2 of their papers we have counts for
5 papers
All in One: Exploring Unified Video-Language Pre-training
Alex Jinpeng Wang, Yixiao Ge, Rui Yan +7
Mainstream Video-Language Pre-training models \cite{actbert,clipbert,violet} consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They purs…
Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation
Shuhe Wang, Yuxian Meng, Xiaofei Sun +5
Multi-modal dialog modeling is of growing interest. In this work, we propose frameworks to resolve a specific case of multi-modal dialog generation that better mimics multi-modal d…
OpenViDial: A Large-Scale, Open-Domain Dialogue Dataset with Visual Contexts
Yuxian Meng, Shuhe Wang, Qinghong Han +4
When humans converse, what a speaker will say next significantly depends on what he sees. Unfortunately, existing dialogue models generate dialogue utterances only based on precedi…
Learning to Customize Model Structures for Few-shot Dialogue Generation Tasks
Yiping Song, Zequn Liu, Wei Bi +2
Training the generative models with minimal corpus is one of the critical challenges for building open-domain dialogue systems. Existing methods tend to use the meta-learning frame…
StalemateBreaker: A Proactive Content-Introducing Approach to Automatic Human-Computer Conversation
Xiang Li, Lili Mou, Rui Yan +1
Existing open-domain human-computer conversation systems are typically passive: they either synthesize or retrieve a reply provided a human-issued utterance. It is generally presum…