6 citations · 6 across the 2 of their papers we have counts for
3 papers
NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
Jagadeesh Balam, Travis Bartley, Edresson Casanova +46
We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder an…
A frontend-backend architecture for tool calls in full-duplex speech models
Ke Hu, Slyne Deng, Chen Chen +10
Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent…
Improving Noise Robustness of an End-to-End Neural Model for Automatic Speech Recognition
Jagadeesh Balam, Jocelyn Huang, Vitaly Lavrukhin +3
We present our experiments in training robust to noise an end-to-end automatic speech recognition (ASR) model using intensive data augmentation. We explore the efficacy of fine-tun…