activity
20182022
most citedLhotse: a speech data representation library for the modern deep learning ecosystem

10 citations · 15 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2021

Beyond Isolated Utterances: Conversational Emotion Recognition

Raghavendra Pappagari, Piotr Żelasko, Jesús Villalba +2

Speech emotion recognition is the task of recognizing the speaker's emotional state given a recording of their utterance. While most of the current approaches focus on inferring em…

cs.CL2021

Joint prediction of truecasing and punctuation for conversational speech in low-resource scenarios

Raghavendra Pappagari, Piotr Żelasko, Agnieszka Mikołajczyk +2

Capitalization and punctuation are important cues for comprehending written texts and conversational transcripts. Yet, many ASR systems do not produce punctuated and case-formatted…

cs.CL2021

What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act Recognition

Piotr Żelasko, Raghavendra Pappagari, Najim Dehak

Dialog acts can be interpreted as the atomic units of a conversation, more fine-grained than utterances, characterized by a specific communicative function. The ability to structur…

cs.CL2020

WER we are and WER we think we are

Piotr Szymański, Piotr Żelasko, Mikolaj Morzy +6

Natural language processing of conversational speech requires the availability of high-quality transcripts. In this paper, we express our skepticism towards the recent reports of v…

cs.CL2020

Punctuation Prediction in Spontaneous Conversations: Can We Mitigate ASR Errors with Retrofitted Word Embeddings?

Łukasz Augustyniak, Piotr Szymanski, Mikołaj Morzy +5

Automatic Speech Recognition (ASR) systems introduce word errors, which often confuse punctuation prediction models, turning punctuation restoration into a challenging task. These…

cs.CL2019

Hierarchical Transformers for Long Document Classification

Raghavendra Pappagari, Piotr Żelasko, Jesús Villalba +2

BERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm. We…