3 papers
cs.SD2026
Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs
Jing Peng, Zichao Nie, Zhisheng Zhang +2
Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated.…
eess.AS2026
LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
Zhisheng Zhang, Xiang Li, Yixuan Zhou +3
Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details…
cs.SD2024
ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models
Zixing Zhang, Weixiang Xu, Zhongren Dong +5
Computational paralinguistics (ComParal) aims to develop algorithms and models to automatically detect, analyze, and interpret non-verbal information from speech communication, e.…