1 paper
Koki Maeda, Tosho Hirasawa, Atsushi Hashimoto +4
Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works o…