1 paper
Tonmoay Deb, Akib Sadmanee, Kishor Kumar Bhaumik +3
While describing Spatio-temporal events in natural language, video captioning models mostly rely on the encoder's latent visual representation. Recent progress on the encoder-decod…