1 paper · 1 filter
Koki Maeda, Tosho Hirasawa, Atsushi Hashimoto +4
Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works o…