8 citations · 22 across the 29 of their papers we have counts for
4 papers · 1 filter
MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
Shaden Shaar, Bradon Thymes, Sirawut Chaixanien +2
Understanding real-world videos such as movies requires integrating visual and dialogue cues. Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and, giv…
Fashionpedia-Ads: Do Your Favorite Advertisements Reveal Your Fashion Taste?
Mengyun Shi, Claire Cardie, Serge Belongie
Consumers are exposed to advertisements across many different domains on the internet, such as fashion, beauty, car, food, and others. On the other hand, fashion represents second…
Fashionpedia-Taste: A Dataset towards Explaining Human Fashion Taste
Mengyun Shi, Serge Belongie, Claire Cardie
Existing fashion datasets do not consider the multi-facts that cause a consumer to like or dislike a fashion image. Even two consumers like a same fashion image, they could like th…
Exploring Visual Engagement Signals for Representation Learning
Menglin Jia, Zuxuan Wu, Austin Reiter +3
Visual engagement in social media platforms comprises interactions with photo posts including comments, shares, and likes. In this paper, we leverage such visual engagement clues a…