2 papers
cs.SD2026
Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models
Yuxuan Chen, Haoyuan Yu, Peize He
Recent spatial self supervised audio models achieve high performance on localization tasks, raising questions about their encoding of microsecond interaural phase fine structures.…
cs.CL2025
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
Yuxuan Chen, Haoyuan Yu
True Full-Duplex (TFD) voice communication--enabling simultaneous listening and speaking with natural turn-taking, overlapping speech, and interruptions--represents a critical mile…