9 papers
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
Séverin Baroudi, Yanis Labrak, Shashi Kumar +7
Extracting patient medical conditions from code-switched clinical spoken dialogues is challenging due to rapid turn-taking and highly overlapped speech. We present a robust system…
Unifying Global and Near-Context Biasing in a Single Trie Pass
Iuliia Thorbecke, Esaú Villatoro-Tello, Juan Zuluaga-Gomez +9
Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and i…
Dialog2Flow: Pre-training Soft-Contrastive Action-Driven Sentence Embeddings for Automatic Dialog Flow Extraction
Sergio Burdisso, Srikanth Madikeri, Petr Motlicek
Efficiently deriving structured workflows from unannotated dialogs remains an underexplored and formidable challenge in computational linguistics. Automating this process could sig…
Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions
Dairazalia Sánchez-Cortés, Sergio Burdisso, Esaú Villatoro-Tello +1
Bias assessment of news sources is paramount for professionals, organizations, and researchers who rely on truthful evidence for information gathering and reporting. While certain…
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +6
In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent pr…
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +5
Self-supervised pretrained models exhibit competitive performance in automatic speech recognition on finetuning, even with limited in-domain supervised data. However, popular pretr…