43 citations · 75 across the 29 of their papers we have counts for
1 paper · 1 filter
Kate Sanders, Benjamin Van Durme
While existing video benchmarks largely consider specialized downstream tasks like retrieval or question-answering (QA), contemporary multimodal AI systems must be capable of well-…