4 papers
Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition
Raphaël Bagat, Zhe Zhang, Junichi Yamagishi +2
Automatic Speech Recognition (ASR) systems, despite achieving remarkable accuracy in general-purpose domains with native speech (L1), struggle in domains like Air Traffic Control (…
VoxEffects: A Speech-Oriented Audio Effects Dataset and Benchmark
Zhe Zhang, Yigitcan Ãzer, Junichi Yamagishi
Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systemat…
The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results
Xingyu Qiu, Yuqian Fu, Jiawei Geng +70
Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across…
A Cross-Perspective Annotated Dataset for Dynamic Object-Level Attention Modeling in Cloud Gaming
Hongqin Lei, Haowei Tang, Zhe Zhang
Cloud gaming has gained popularity as it provides high-quality gaming experiences on thin hardware, such as phones and tablets. Transmitting gameplay frames at high resolutions and…