1 citations · 1 across the 12 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.SD2026
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar +6
Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR aro…
cs.SD2026
Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar +8
An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model wri…