5 papers
CooT: Learning to Coordinate In-Context with Coordination Transformers
Huai-Chih Wang, Hsiang-Chun Chuang, Hsi-Chun Cheng +2
Effective coordination among unfamiliar partners remains a major challenge in multi-agent systems. Existing approaches, such as population-based methods, improve robustness through…
An Exploration of Mamba for Speech Self-Supervised Models
Tzu-Quan Lin, Heng-Cheng Kuo, Tzu-Chieh Wei +5
While Mamba has demonstrated strong performance in language modeling, its potential as a speech self-supervised learning (SSL) model remains underexplored, with prior studies limit…
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
Tzu-Quan Lin, Hsi-Chun Cheng, Hung-yi Lee +1
In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications. However, little research has explored how these models encode speak…
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
Cheng-Kang Chou, Chan-Jan Hsu, Ho-Lam Chung +5
We propose a self-refining framework that enhances ASR performance with only unlabeled datasets. The process starts with an existing ASR model generating pseudo-labels on unannotat…
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang +77
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spo…