1 paper
Reza Ghoddoosian, Nakul Agarwal, Isht Dwivedi +1
Vision-language models (VLMs) are capable of recognizing unseen actions. However, existing VLMs lack intrinsic understanding of procedural action concepts. Hence, they overfit to f…