4 papers
Unifying Model and Layer Fusion for Speech Foundation Models
Yi-Jen Shih, David Harwath
Speech Foundation Models have gained significant attention recently. Prior works have shown that the fusion of representations from multiple layers of the same model or the fusion…
Can Speech LLMs Think while Listening?
Yi-Jen Shih, Desh Raj, Chunyang Wu +6
Recent advances in speech large language models (speech LLMs) have enabled seamless spoken interactions, but these systems still struggle with complex reasoning tasks. Previously,…
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang +77
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spo…
Measuring Sound Symbolism in Audio-visual Models
Wei-Cheng Tseng, Yi-Jen Shih, David Harwath +1
Audio-visual pre-trained models have gained substantial attention recently and demonstrated superior performance on various audio-visual tasks. This study investigates whether pre-…