Auxiliary Cross-Modal Representation Learning with Triplet Loss Functions for Online Handwriting Recognition
arXiv:2202.07901 · doi:10.1109/ACCESS.2023.3310819
Abstract
Cross-modal representation learning learns a shared embedding between two or more modalities to improve performance in a given task compared to using only one of the modalities. Cross-modal representation learning from different data types -- such as images and time-series data (e.g., audio or text data) -- requires a deep metric learning loss that minimizes the distance between the modality embeddings. In this paper, we propose to use the contrastive or triplet loss, which uses positive and negative identities to create sample pairs with different labels, for cross-modal representation learning between image and time-series modalities (CMR-IS). By adapting the triplet loss for cross-modal representation learning, higher accuracy in the main (time-series classification) task can be achieved by exploiting additional information of the auxiliary (image classification) task. We present a triplet loss with a dynamic margin for single label and sequence-to-sequence classification tasks. We perform extensive evaluations on synthetic image and time-series data, and on data for offline handwriting recognition (HWR) and on online HWR from sensor-enhanced pens for classifying written words. Our experiments show an improved classification accuracy, faster convergence, and better generalizability due to an improved cross-modal representation. Furthermore, the more suitable generalizability leads to a better adaptability between writers for online HWR.
References in corpus (16)
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- COCOA: Cross Modality Contrastive Learning for Sensor Data
- Full Page Handwriting Recognition via Image to Sequence Extraction
- Content and Style Aware Generation of Text-line Images for Handwriting Recognition
- Domain Adaptation for Time-Series Classification to Mitigate Covariate Shift
- CrossATNet - A Novel Cross-Attention Based Framework for Sketch-Based Image Retrieval
- Benchmarking Online Sequence-to-Sequence and Character-based Handwriting Recognition from IMU-Enhanced Pens
- A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval
- Beyond Just Vision: A Review on Self-Supervised Representation Learning on Multimodal and Temporal Data
- Multimodal Self-Supervised Learning of General Audio Representations
- ConceptBeam: Concept Driven Target Speech Extraction
- MM-ALT: A Multimodal Automatic Lyric Transcription System
- Improving Accuracy and Explainability of Online Handwriting Recognition
- Representation Learning for Tablet and Paper Domain Adaptation in Favor of Online Handwriting Recognition
- Motion-Based Handwriting Recognition
- Detecting Handwritten Mathematical Terms with Sensor Based Data