4 papers
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Tianhong Zhou, Mingyang Han, Boyu Li +8
Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models ex…
FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing
Yuxuan Jiang, Mingyang Han, Yusheng Dai +12
Text-to-audio (TTA) generation has made significant strides, yet achieving precise and consistent audio editing remains a major challenge. However, existing methods struggle to bal…
VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding
Jashin Ye, Dongxiao Wang, Yixuan Ye +10
While large audio language models (LALMs) have achieved remarkable progress in audio processing at the second- or minute-level scale, understanding hour-level audio remains a funda…
Improving Large-scale Deep Biasing with Phoneme Features and Text-only Data in Streaming Transducer
Jin Qiu, Lu Huang, Boyu Li +3
Deep biasing for the Transducer can improve the recognition performance of rare words or contextual entities, which is essential in practical applications, especially for streaming…