From the 1 of 7 linked papers with an AI index.
7 papers
Voice Memory for Agentic Speech Recognition
Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko +3
The paper introduces Voice Memory, an inference-only framework for agentic speech recognition that uses a frozen corrector and a per‑domain memory file to decide when to modify hyp…
Test-Time Alignment for Large Language Models via Textual Model Predictive Control
Kuang-Da Wang, Teng-Ruei Chen, Yu Heng Hung +7
Aligning Large Language Models (LLMs) with human preferences through finetuning is resource-intensive, motivating lightweight alternatives at test time. We address test-time alignm…
SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
Chih-Kai Yang, Yen-Ting Piao, Tzu-Wen Hsu +8
Knowledge editing enables targeted updates without retraining, but prior work focuses on textual or visual facts, leaving abstract auditory perceptual knowledge underexplored. We i…
NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
Yen-Ting Lin, Zhehuai Chen, Piotr Zelasko +11
Construction of a general-purpose post-recognition error corrector poses a crucial question: how can we most effectively train a model on a large mixture of domain datasets? The an…
Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
Bo-Han Feng, Chien-Feng Liu, Yu-Hsuan Li Liang +9
Large audio-language models (LALMs) extend text-based LLMs with auditory understanding, offering new opportunities for multimodal applications. While their perception, reasoning, a…
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang +77
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spo…