YFCC100M: The New Data in Multimedia Research
arXiv:1503.01817 · doi:10.1145/2812802
Abstract
We present the Yahoo Flickr Creative Commons 100 Million Dataset (YFCC100M), the largest public multimedia collection that has ever been released. The dataset contains a total of 100 million media objects, of which approximately 99.2 million are photos and 0.8 million are videos, all of which carry a Creative Commons license. Each media object in the dataset is represented by several pieces of metadata, e.g. Flickr identifier, owner name, camera, title, tags, geo, media source. The collection provides a comprehensive snapshot of how photos and videos were taken, described, and shared over the years, from the inception of Flickr in 2004 until early 2014. In this article we explain the rationale behind its creation, as well as the implications the dataset has for science, research, engineering, and development. We further present several new challenges in multimedia research that can now be expanded upon with our dataset.
References in corpus (2)
Cited by in corpus (232)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Learning Transferable Visual Models From Natural Language Supervision
- Zero-Shot Text-to-Image Generation
- Momentum Contrast for Unsupervised Visual Representation Learning
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- Generating Videos with Scene Dynamics
- KonIQ-10k: Towards an ecologically valid and large-scale IQA database
- KonIQ-10k: An ecologically valid database for deep learning of blind image quality assessment
- Reproducible scaling laws for contrastive language-image learning
- YOLO9000: Better, Faster, Stronger
- Self-Supervised Representation Learning: Introduction, Advances and Challenges
- Deep Learning is Robust to Massive Label Noise
- Billion-scale semi-supervised learning for image classification
- Image Matching across Wide Baselines: From Paper to Practice
- A Review on Deep Learning Techniques for Video Prediction
- Cross-Age LFW: A Database for Studying Cross-Age Face Recognition in Unconstrained Environments
- Image Quality Assessment using Contrastive Learning
- Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder
- CosFace: Large Margin Cosine Loss for Deep Face Recognition
- Visual Question Answering: Datasets, Algorithms, and Future Challenges
- Self-training with Noisy Student improves ImageNet classification
- Detecting Sarcasm in Multimodal Social Platforms
- XCiT: Cross-Covariance Image Transformers
- SoundNet: Learning Sound Representations from Unlabeled Video
- Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- Patch-VQ: 'Patching Up' the Video Quality Problem
- Few-Example Object Detection with Model Communication
- Medical Visual Question Answering: A Survey
- On Low-Resolution Face Recognition in the Wild: Comparisons and New Techniques
- Learning a Text-Video Embedding from Incomplete and Heterogeneous Data
- GraphIQA: Learning Distortion Graph Representations for Blind Image Quality Assessment
- A Survey on Long-Tailed Visual Recognition
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- Multimodal Data Integration for Oncology in the Era of Deep Neural Networks: A Review
- EmbraceNet: A robust deep learning architecture for multimodal classification
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
- Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset Biases
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- Co-learning: Learning from Noisy Labels with Self-supervision
- Real-Time Adaptive Image Compression
- RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
- ARBEE: Towards Automated Recognition of Bodily Expression of Emotion In the Wild
- Causally Regularized Learning with Agnostic Data Selection Bias
- Fluency-Guided Cross-Lingual Image Captioning
- Extracting human emotions at different places based on facial expressions and spatial clustering analysis
- Local Feature Matching Using Deep Learning: A Survey
- Unsupervised Sound Separation Using Mixture Invariant Training
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- Domain Adaptive Transfer Learning with Specialist Models
- Unsupervised Feature Learning Based on Deep Models for Environmental Audio Tagging
- Learning Points and Routes to Recommend Trajectories
- Deep Learning for Video Classification and Captioning
- Learning from Noisy Labels with Distillation
- Identifying and Compensating for Feature Deviation in Imbalanced Deep Learning
- Learning Two-View Correspondences and Geometry Using Order-Aware Network
- Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted Approach
- Embracing Error to Enable Rapid Crowdsourcing
- Few-Shot Bot: Prompt-Based Learning for Dialogue Systems
- CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval
- Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision People
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds
- Gender and Racial Bias in Visual Question Answering Datasets
- Support-set bottlenecks for video-text representation learning
- MultiGrain: a unified image embedding for classes and instances
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- Radioactive data: tracing through training
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- Visual Chirality
- AdaCos: Adaptively Scaling Cosine Logits for Effectively Learning Deep Face Representations
- On Model Calibration for Long-Tailed Object Detection and Instance Segmentation
- A Survey on Point-of-Interest Recommendations Leveraging Heterogeneous Data
- Partially-Supervised Image Captioning
- Deep CNN Framework for Audio Event Recognition using Weakly Labeled Web Data
- Rethinking Benchmarks for Cross-modal Image-text Retrieval
- WebFace260M: A Benchmark Unveiling the Power of Million-Scale Deep Face Recognition
- MemexQA: Visual Memex Question Answering
- ContextDesc: Local Descriptor Augmentation with Cross-Modality Context
- Towards Non-I.I.D. Image Classification: A Dataset and Baselines
- Significantly improving zero-shot X-ray pathology classification via fine-tuning pre-trained image-text encoders
- Certified Data Removal from Machine Learning Models
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Analysis and Optimization of fastText Linear Text Classifier
- A Simple Fine-tuning Is All You Need: Towards Robust Deep Learning Via Adversarial Fine-tuning
- Picture It In Your Mind: Generating High Level Visual Representations From Textual Descriptions
- Augmenting Transformers with KNN-Based Composite Memory for Dialogue
- Comparison of State-of-the-Art Deep Learning APIs for Image Multi-Label Classification using Semantic Metrics
- Video Relation Detection via Tracklet based Visual Transformer
- In Defense of Grid Features for Visual Question Answering
- Repurposing Existing Deep Networks for Caption and Aesthetic-Guided Image Cropping
- How to Train Your MAML to Excel in Few-Shot Classification
- Health-promoting Potential of Parks in 35 Cities Worldwide
- Deep Multi-View Stereo gone wild
- Content-Aware Detection of Temporal Metadata Manipulation
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- Breaking the Trilemma of Privacy, Utility, Efficiency via Controllable Machine Unlearning
- Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
- Transfer Learning From Sound Representations For Anger Detection in Speech
- Simultaneous Learning of Trees and Representations for Extreme Classification and Density Estimation
- TopicFM+: Boosting Accuracy and Efficiency of Topic-Assisted Feature Matching
- DC3DCD: unsupervised learning for multiclass 3D point cloud change detection
- Improving Video Generation for Multi-functional Applications
- A Secure Mobile Authentication Alternative to Biometrics
- Visual Relationship Detection using Scene Graphs: A Survey
- OG-SGG: Ontology-Guided Scene Graph Generation. A Case Study in Transfer Learning for Telepresence Robotics
- Improving Borderline Adulthood Facial Age Estimation through Ensemble Learning
- MLM: A Benchmark Dataset for Multitask Learning with Multiple Languages and Modalities
- Learning to Guide Local Feature Matches
- Exploring Masked Autoencoders for Sensor-Agnostic Image Retrieval in Remote Sensing
- Beautiful and damned. Combined effect of content quality and social ties on user engagement
- Fine-grained Video Attractiveness Prediction Using Multimodal Deep Learning on a Large Real-world Dataset
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- Blind Predicting Similar Quality Map for Image Quality Assessment
- INTERN: A New Learning Paradigm Towards General Vision
- Changing Fashion Cultures
- Learnable Motion Coherence for Correspondence Pruning
- Audio-Visual Quality Assessment for User Generated Content: Database and Method
- CCMB: A Large-scale Chinese Cross-modal Benchmark
- Exploring AI in Fashion: A Review of Aesthetics, Personalization, Virtual Try-On, and Forecasting
- Progressive Correspondence Pruning by Consensus Learning
- Déjà Vu: an empirical evaluation of the memorization properties of ConvNets
- Large Scale Clustering with Variational EM for Gaussian Mixture Models
- S2DNet: Learning Accurate Correspondences for Sparse-to-Dense Feature Matching
- Flickr Africa: Examining Geo-Diversity in Large-Scale, Human-Centric Visual Data
- Interactive Video Corpus Moment Retrieval using Reinforcement Learning
- Learning Probabilistic Coordinate Fields for Robust Correspondences
- Dual-Glance Model for Deciphering Social Relationships
- Selective Deep Convolutional Features for Image Retrieval
- From Selective Deep Convolutional Features to Compact Binary Representations for Image Retrieval
- Attention and Localization based on a Deep Convolutional Recurrent Model for Weakly Supervised Audio Tagging
- AdaLAM: Revisiting Handcrafted Outlier Detection
- A2B: Anchor to Barycentric Coordinate for Robust Correspondence
- Geo-Aware Networks for Fine-Grained Recognition
- ClusterFit: Improving Generalization of Visual Representations
- Similarity Join and Similarity Self-Join Size Estimation in a Streaming Environment
- Multi-Label Meta Weighting for Long-Tailed Dynamic Scene Graph Generation
- Tag Prediction at Flickr: a View from the Darkroom
- A Hybrid Filtering for Micro-video Hashtag Recommendation using Graph-based Deep Neural Network
- Learning without Prejudice: Avoiding Bias in Webly-Supervised Action Recognition
- ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG
- Probabilistic Video Generation using Holistic Attribute Control
- The network structure of visited locations according to geotagged social media photos
- Happy Travelers Take Big Pictures: A Psychological Study with Machine Learning and Big Data
- Powers of layers for image-to-image translation
- The 2021 Image Similarity Dataset and Challenge
- nSimplex Zen: A Novel Dimensionality Reduction for Euclidean and Hilbert Spaces
- Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization
- Online Continual Learning with Natural Distribution Shifts: An Empirical Study with Visual Data
- RANSAC-Flow: generic two-stage image alignment
- Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps
- High-contrast "gaudy" images improve the training of deep neural network models of visual cortex
- Train and You'll Miss It: Interactive Model Iteration with Weak Supervision and Pre-Trained Embeddings
- DeeSIL: Deep-Shallow Incremental Learning
- Event Detection and Retrieval on Social Media
- Image recognition from raw labels collected without annotators
- Are CLIP features all you need for Universal Synthetic Image Origin Attribution?
- A Jointly Learned Context-Aware Place of Interest Embedding for Trip Recommendations
- Channel-Level Variable Quantization Network for Deep Image Compression
- Multi-modal Geolocation Estimation Using Deep Neural Networks
- Simplicity Bias Leads to Amplified Performance Disparities
- Towards Automatic Construction of Diverse, High-quality Image Dataset
- Improving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attention
- Recommending POIs for Tourists by User Behavior Modeling and Pseudo-Rating
- Exquisitor: Interactive Learning at Large
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval
- Enriching Ontology with Temporal Commonsense for Low-Resource Audio Tagging
- FooDI-ML: a large multi-language dataset of food, drinks and groceries images and descriptions
- MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object Detection
- LOH and behold: Web-scale visual search, recommendation and clustering using Locally Optimized Hashing
- MUSIQ: Multi-scale Image Quality Transformer
- Limited Gradient Descent: Learning With Noisy Labels
- V3C - a Research Video Collection
- Video Relationship Detection Using Mixture of Experts
- OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
- Reinforcement Learning for Strategic Recommendations
- Revisiting IM2GPS in the Deep Learning Era
- Unsupervised Discriminative Learning of Sounds for Audio Event Classification
- Leveraging Multiple Relations for Fashion Trend Forecasting Based on Social Media
- DCAR: A Discriminative and Compact Audio Representation to Improve Event Detection
- LLHA-Net: A Hierarchical Attention Network for Two-View Correspondence Learning
- OSOA: One-Shot Online Adaptation of Deep Generative Models for Lossless Compression
- Generalized Few-Shot Video Classification with Video Retrieval and Feature Generation
- MVImgNet2.0: A Larger-scale Dataset of Multi-view Images
- A Link Between the Multiplicative and Additive Functional Asplund's Metrics
- Practical Near Neighbor Search via Group Testing
- Self-Distilled Self-Supervised Representation Learning
- Automatic Dataset Augmentation
- Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality Assessment
- Chameleon: A Semi-AutoML framework targeting quick and scalable development and deployment of production-ready ML systems for SMEs
- Graph convolutional networks for learning with few clean and many noisy labels
- 3D-FCT: Simultaneous 3D Object Detection and Tracking Using Feature Correlation
- Digital Collections Explorer: An Open-Source, Multimodal Viewer for Searching Digital Collections
- Fast k-means based on KNN Graph
- An evaluation of large-scale methods for image instance and class discovery
- ZoDIAC: Zoneout Dropout Injection Attention Calculation
- Clustering via Boundary Erosion
- Towards Generative Location Awareness for Disaster Response: A Probabilistic Cross-view Geolocalization Approach
- Leaking Sensitive Financial Accounting Data in Plain Sight using Deep Autoencoder Neural Networks
- Noise-Tolerant Hybrid Prototypical Learning with Noisy Web Data
- Video Stream Retrieval of Unseen Queries using Semantic Memory
- OrigamiSet1.0: Two New Datasets for Origami Classification and Difficulty Estimation
- Learning to Localize Sound Sources in Visual Scenes: Analysis and Applications
- Divergence Optimization for Noisy Universal Domain Adaptation
- Localizing Visual Sounds the Hard Way
- Weakly Supervised Dataset Collection for Robust Person Detection
- ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation
- Stochastic Neighbor Embedding of Multimodal Relational Data for Image-Text Simultaneous Visualization
- Fairness Properties of Face Recognition and Obfuscation Systems
- Point of Interest Recommendation: Pitfalls and Viable Solutions
- Automatic Photo to Ideophone Manga Matching
- Procrustean Training for Imbalanced Deep Learning
- From Photo Streams to Evolving Situations
- Tackling the Problem of Limited Data and Annotations in Semantic Segmentation
- Fit to Measure: Reasoning about Sizes for Robust Object Recognition
- Real-time Spatio-temporal Event Detection on Geotagged Social Media
- Does Face Recognition Error Echo Gender Classification Error?
- Learning to Generate Novel Scene Compositions from Single Images and Videos
- A Classification approach towards Unsupervised Learning of Visual Representations
- Towards Learning Cross-Modal Perception-Trace Models
- User-Aware Folk Popularity Rank: User-Popularity-Based Tag Recommendation That Can Enhance Social Popularity
- What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering
- Self-Supervised Visual Representations Learning by Contrastive Mask Prediction
- What Can Style Transfer and Paintings Do For Model Robustness?
- GeoWINE: Geolocation based Wiki, Image,News and Event Retrieval
- A Self-Explainable Stylish Image Captioning Framework via Multi-References
- Structured Recommendation
- A Practical Guide to Streaming Continual Learning
- Speeding up the Köhler's method of contrast thresholding
- MIML-FCN+: Multi-instance Multi-label Learning via Fully Convolutional Networks with Privileged Information
- VideoMCC: a New Benchmark for Video Comprehension
- Learning a Dynamic Map of Visual Appearance
- PixelTransformer: Sample Conditioned Signal Generation