Voice Recognition Algorithms using Mel Frequency Cepstral Coefficient (MFCC) and Dynamic Time Warping (DTW) Techniques
arXiv:1003.4083
Abstract
Digital processing of speech signal and voice recognition algorithm is very important for fast and accurate automatic voice recognition technology. The voice is a signal of infinite information. A direct analysis and synthesizing the complex voice signal is due to too much information contained in the signal. Therefore the digital signal processes such as Feature Extraction and Feature Matching are introduced to represent the voice signal. Several methods such as Liner Predictive Predictive Coding (LPC), Hidden Markov Model (HMM), Artificial Neural Network (ANN) and etc are evaluated with a view to identify a straight forward and effective method for voice signal. The extraction and matching process is implemented right after the Pre Processing or filtering signal is performed. The non-parametric method for modelling the human auditory perception system, Mel Frequency Cepstral Coefficients (MFCCs) are utilize as extraction techniques. The non linear sequence alignment known as Dynamic Time Warping (DTW) introduced by Sakoe Chiba has been used as features matching techniques. Since it's obvious that the voice signal tends to have different temporal rate, the alignment is important to produce the better performance.This paper present the viability of MFCC to extract features and DTW to compare the test patterns.
Cited by in corpus (27)
- In-situ process monitoring and adaptive quality enhancement in laser additive manufacturing: a critical review
- In-situ crack and keyhole pore detection in laser directed energy deposition through acoustic signal and deep learning
- Principal Component Analysis-Linear Discriminant Analysis Feature Extractor for Pattern Recognition
- Security and Privacy in the Emerging Cyber-Physical World: A Survey
- Concurrent Activity Recognition with Multimodal CNN-LSTM Structure
- Improving Zero-shot Voice Style Transfer via Disentangled Representation Learning
- Thank you for Attention: A survey on Attention-based Artificial Neural Networks for Automatic Speech Recognition
- A Multi-Biometrics for Twins Identification Based Speech and Ear
- An Extensive Analysis of Query by Singing/Humming System Through Query Proportion
- Masked Pre-trained Encoder base on Joint CTC-Transformer
- Deep Neural Network Based Respiratory Pathology Classification Using Cough Sounds
- Cross-modal Variational Auto-encoder with Distributed Latent Spaces and Associators
- An Attention-Based Speaker Naming Method for Online Adaptation in Non-Fixed Scenarios
- Exploiting Fully Convolutional Network and Visualization Techniques on Spontaneous Speech for Dementia Detection
- Asymmetric Learning Vector Quantization for Efficient Nearest Neighbor Classification in Dynamic Time Warping Spaces
- Multimodal Systems: Taxonomy, Methods, and Challenges
- LSTM and GPT-2 Synthetic Speech Transfer Learning for Speaker Recognition to Overcome Data Scarcity
- Modelling Animal Biodiversity Using Acoustic Monitoring and Deep Learning
- Using Deep Learning Techniques and Inferential Speech Statistics for AI Synthesised Speech Recognition
- Audio Adversarial Examples: Attacks Using Vocal Masks
- WaDeNet: Wavelet Decomposition based CNN for Speech Processing
- Towards robust audio spoofing detection: a detailed comparison of traditional and learned features
- LaNet: Real-time Lane Identification by Learning Road SurfaceCharacteristics from Accelerometer Data
- APB2Face: Audio-guided face reenactment with auxiliary pose and blink signals
- Contactless Cardiac Arrest Detection Using Smart Devices
- Smart Speakers, the Next Frontier in Computational Health
- Trainable Time Warping: Aligning Time-Series in the Continuous-Time Domain