86 citations · 86 across the 3 of their papers we have counts for
3 papers
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
Vahid Noroozi, Zhehuai Chen, Somshubra Majumdar +3
In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. A…
Investigating End-to-End ASR Architectures for Long Form Audio Transcription
Nithin Rao Koluguri, Samuel Kriman, Georgy Zelenfroind +5
This paper presents an overview and evaluation of some of the end-to-end ASR models on long-form audios. We study three categories of Automatic Speech Recognition(ASR) models based…
Joint Deep Modeling of Users and Items Using Reviews for Recommendation
Lei Zheng, Vahid Noroozi, Philip S. Yu
A large amount of information exists in reviews written by users. This source of information has been ignored by most of the current recommender systems while it can potentially al…