1 paper
HyoJung Han, Mohamed Anwar, Juan Pino +4
Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potent…