activity
20182026
most citedDeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

15 citations · 27 across the 36 of their papers we have counts for

collaborators
Showing cs.CLShow all

19 papers · 1 filter

cs.CL2026

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

Jagadeesh Balam, Travis Bartley, Edresson Casanova +46

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder an…

cs.CL2026

A frontend-backend architecture for tool calls in full-duplex speech models

Ke Hu, Slyne Deng, Chen Chen +10

Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent…

cs.CL2026

Voice Memory for Agentic Speech Recognition

Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko +3

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance wh…

cs.CL2025★ 1 cited

SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models

Zhen Wan, Chao-Han Huck Yang, Yahan Yu +8

We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models, LLM Voice, design…

cs.CL2025

Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Ke Hu, Krishna Puvvada, Elena Rastorgueva +7

We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks su…

cs.CL2025

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Ke Hu, Ehsan Hosseini-Asl, Chen Chen +7

Spoken dialogue is an intuitive form of human-computer interaction, yet current speech language models often remain constrained to turn-based exchanges, lacking real-time adaptabil…