paper

PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio

arXiv:2603.05128

Abstract

Large Audio Language Models (LALMs) are increasingly capable of reasoning over audio, yet existing benchmarks offer limited coverage of reasoning in polyphonic audio, where multiple sound events co-occur and induce compositional structure. To address this gap, we introduce PolyBench, a benchmark designed to evaluate compositional reasoning in polyphonic audio, comprising five evaluation subsets that cover counting, classification, detection, concurrency, and duration estimation, all of which require reasoning over multiple concurrent events and their relations. Our evaluation of state-of-the-art LALMs reveals consistent performance degradation in polyphonic settings, indicating a fundamental bottleneck in current LALMs.

Accepted by INTERSPEECH 2026

PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio · wovepaper