PFML: Self-Supervised Learning of Time-Series Data Without Representation Collapse
arXiv:2411.10087 · doi:10.1109/ACCESS.2025.3556957
Abstract
Self-supervised learning (SSL) is a data-driven learning approach that utilizes the innate structure of the data to guide the learning process. In contrast to supervised learning, which depends on external labels, SSL utilizes the inherent characteristics of the data to produce its own supervisory signal. However, one frequent issue with SSL methods is representation collapse, where the model outputs a constant input-invariant feature representation. This issue hinders the potential application of SSL methods to new data modalities, as trying to avoid representation collapse wastes researchers' time and effort. This paper introduces a novel SSL algorithm for time-series data called Prediction of Functionals from Masked Latents (PFML). Instead of predicting masked input signals or their latent representations directly, PFML operates by predicting statistical functionals of the input signal corresponding to masked embeddings, given a sequence of unmasked embeddings. The algorithm is designed to avoid representation collapse, rendering it straightforwardly applicable to different time-series data domains, such as novel sensor modalities in clinical data. We demonstrate the effectiveness of PFML through complex, real-life classification tasks across three different data modalities: infant posture and movement classification from multi-sensor inertial measurement unit data, emotion recognition from speech data, and sleep stage classification from EEG data. The results show that PFML is superior to a conceptually similar SSL method and a contrastive learning-based SSL method. Additionally, PFML is on par with the current state-of-the-art SSL method, while also being conceptually simpler and without suffering from representation collapse.
Accepted for publication in IEEE Access
References in corpus (5)
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- Understanding Dimensional Collapse in Contrastive Self-supervised Learning
- Evaluation of self-supervised pre-training for automatic infant movement classification using wearable movement sensors