3 papers
cs.CL2026
Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
Ryan Soh-Eun Shim, Domenico De Cristofaro, Chengzhi Martin Hu +2
Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on sp…
cs.CL2026
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
Domenico De Cristofaro, Alessandro Vietti, Marianne Pouplier +1
Recent studies have shown that intermediate layers in multilingual speech models often encode more phonetically accurate representations than the final output layer. In this work,…
cs.CL2025
Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
Domenico De Cristofaro, Vincenzo Norman Vitale, Alessandro Vietti
Automatic Speech Recognition has advanced with self-supervised learning, enabling feature extraction directly from raw audio. In Wav2Vec, a CNN first transforms audio into feature…