activity
20192026
most citedSpeech Emotion Recognition using Self-Supervised Features

5 citations · 8 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark

Arnon Turetzky, Avihu Dekel, Hagai Aronowitz +2

Spoken meaning often depends not only on what is said, but also on which word is emphasized. The same sentence can convey correction, contrast, or clarification depending on where…

cs.CL2025

Advancing Speech Understanding in Speech-Aware Language Models with GRPO

Avishai Elmakies, Hagai Aronowitz, Nimrod Shabtay +3

In this paper, we introduce a Group Relative Policy Optimization (GRPO)-based method for training Speech-Aware Large Language Models (SALLMs) on open-format speech understanding ta…

cs.CL2022

Towards a Common Speech Analysis Engine

Hagai Aronowitz, Itai Gat, Edmilson Morais +2

Recent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-super…

cs.CL2022

A new data augmentation method for intent classification enhancement and its application on spoken conversation datasets

Zvi Kons, Aharon Satt, Hong-Kwang Kuo +4

Intent classifiers are vital to the successful operation of virtual agent systems. This is especially so in voice activated systems where the data can be noisy with many ambiguous…

cs.CL2021

RNN Transducer Models For Spoken Language Understanding

Samuel Thomas, Hong-Kwang J. Kuo, George Saon +5

We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding(SLU). These end-to-end (E2E) models are constructed in thr…

cs.CL2020

Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems

Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas +5

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is…