2 papers
eess.AS2026
Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization
Xiangyu Zhang, Benjamin John Southwell, Siqi Pan +3
Audio tokenization has emerged as a critical component in end-to-end audio language models, enabling efficient discrete representation learning for both audio understanding and gen…
eess.AS2025
Sub-band Domain Multi-Hypothesis Acoustic Echo Canceler Based Acoustic Scene Analysis
Benjamin J Southwell, Yin-Lee Ho, David Gunawan
This paper introduces a novel approach for acoustic scene analysis by exploiting an ensemble of statistics extracted from a sub-band domain multi-hypothesis acoustic echo canceler…