6 citations · 7 across the 2 of their papers we have counts for
4 papers
Scaling Multimodal Pre-Training via Cross-Modality Gradient Harmonization
Junru Wu, Yi Liang, Feng Han +3
Self-supervised pre-training recently demonstrates success on large-scale multimodal data, and state-of-the-art contrastive learning methods often enforce the feature consistency f…
Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language
Hassan Akbari, Hamid Palangi, Jianwei Yang +6
Neuro-symbolic representations have proved effective in learning structure information in vision and language. In this paper, we propose a new model architecture for learning multi…
Multi-level Multimodal Common Semantic Space for Image-Phrase Grounding
Hassan Akbari, Svebor Karaman, Surabhi Bhargava +3
We address the problem of phrase grounding by lear ing a multi-level common semantic space shared by the textual and visual modalities. We exploit multiple levels of feature maps o…
Lip2AudSpec: Speech reconstruction from silent lip movements video
Hassan Akbari, Himani Arora, Liangliang Cao +1
In this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos. We use auditory spectrogram as spectral representation of sp…