1 paper
Yolo Y. Tang, Daiki Shimada, Jiayue Meng +14
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models main…