3 papers
cs.CL2026
NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
Jagadeesh Balam, Travis Bartley, Edresson Casanova +46
We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder an…
cs.CL2026
A frontend-backend architecture for tool calls in full-duplex speech models
Ke Hu, Slyne Deng, Chen Chen +10
Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent…
cs.CV2026
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
Amrita Mazumdar, Seonwook Park, Rajarshi Roy +6
Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smile…