315 citations · 331 across the 7 of their papers we have counts for
12 papers
Learning a Grammar Inducer from Massive Uncurated Instructional Videos
Songyang Zhang, Linfeng Song, Lifeng Jin +4
Video-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text. While previous work focuses on building systems…
Make-A-Video: Text-to-Video Generation without Text-Video Data
Uriel Singer, Adam Polyak, Thomas Hayes +10
We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: le…
MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration
Thomas Hayes, Songyang Zhang, Xi Yin +6
Multimodal video-audio-text understanding and generation can benefit from datasets that are narrow but rich. The narrowness allows bite-sized challenges that the research community…
Video-aided Unsupervised Grammar Induction
Songyang Zhang, Linfeng Song, Lifeng Jin +3
We investigate video-aided grammar induction, which learns a constituency parser from both unlabeled text and its corresponding video. Existing methods of multi-modal grammar induc…
SAT: 2D Semantics Assisted Training for 3D Visual Grounding
Zhengyuan Yang, Songyang Zhang, Liwei Wang +1
3D visual grounding aims at grounding a natural language description about a 3D scene, usually represented in the form of 3D point clouds, to the targeted object region. Point clou…
Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language
Songyang Zhang, Houwen Peng, Jianlong Fu +2
We address the problem of retrieving a specific moment from an untrimmed video by natural language. It is a challenging problem because a target moment may take place in the contex…