Transcribing Content from Structural Images with Spotlight Mechanism
arXiv:1905.10954 · doi:10.1145/3219819.3219962
Abstract
Transcribing content from structural images, e.g., writing notes from music scores, is a challenging task as not only the content objects should be recognized, but the internal structure should also be preserved. Existing image recognition methods mainly work on images with simple content (e.g., text lines with characters), but are not capable to identify ones with more complex content (e.g., structured symbols), which often follow a fine-grained grammar. To this end, in this paper, we propose a hierarchical Spotlight Transcribing Network (STN) framework followed by a two-stage "where-to-what" solution. Specifically, we first decide "where-to-look" through a novel spotlight mechanism to focus on different areas of the original image following its structure. Then, we decide "what-to-write" by developing a GRU based network with the spotlight areas for transcribing the content accordingly. Moreover, we propose two implementations on the basis of STN, i.e., STNM and STNR, where the spotlight movement follows the Markov property and Recurrent modeling, respectively. We also design a reinforcement method to refine the framework by self-improving the spotlight mechanism. We conduct extensive experiments on many structural image datasets, where the results clearly demonstrate the effectiveness of STN framework.
Accepted by KDD2018 Research Track. In proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD'18)
References in corpus (10)
- Adam: A Method for Stochastic Optimization
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Neural Machine Translation by Jointly Learning to Align and Translate
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- Recurrent Models of Visual Attention
- Sequence Level Training with Recurrent Neural Networks
- An Actor-Critic Algorithm for Sequence Prediction
- Recursive Recurrent Nets with Attention Modeling for OCR in the Wild
Cited by in corpus (6)
- Exploiting Cognitive Structure for Adaptive Learning
- EKT: Exercise-aware Knowledge Tracing for Student Performance Prediction
- Adam revisited: a weighted past gradients perspective
- QuesNet: A Unified Representation for Heterogeneous Test Questions
- EDSL: An Encoder-Decoder Architecture with Symbol-Level Features for Printed Mathematical Expression Recognition
- Neural Cognitive Diagnosis for Intelligent Education Systems