1 paper
Chunhui Zhang, Yiren Jian, Zhongyu Ouyang +1
Developing video captioning models is computationally expensive. The dynamic nature of video also complicates the design of multimodal models that can effectively caption these seq…