1 paper · 1 filter
Zhenyu Lu, Lakshay Sethi
Previous methods for audio-image matching generally fall into one of two categories: pipeline models or End-to-End models. Pipeline models first transcribe speech and then encode t…