1 paper
Andrew Franck, Brendan Ng, Ben Fitzgerald +3
Video-language benchmarks are usually constructed by the dataset authors without published reliability statistics, leaving the noise floor of the construct unknown. We argue that m…