activity
20242026
most citedAudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation

2 citations · 2 across the 14 of their papers we have counts for

collaborators

18 papers

eess.SP2026

BeamFocusNet: Beamforming-based Explicit Spatial Signal Focusing for Robust DoA Estimation under Low-SNR and Single-Snapshot Conditions

Xuyao Deng, Qisheng Xu, Shuo Liu +3

DoA estimation plays a crucial role in signal processing. Inspired by beamforming, recent works employ neural networks to estimate filters for filtering received signals to achieve…

eess.AS2026

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning

Kele Xu, Yulu Fang, Boda Zhou +6

This paper examines audio self-supervised learning (SSL) through the alignment between pretraining objectives, architectural inductive biases, and downstream applications. Rather t…

cs.SD2026

AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models

Hui Geng, Yi Su, Zijian Gao +5

Recent advances in pretrained large audio-language models (LALMs) have demonstrated strong capabilities across speech, sound, and music. To adapt these models to downstream tasks w…

cs.CV2026

NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

Xin Li, Yeying Jin, Suhang Yao +95

This paper presents an overview of the NTIRE 2026 Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images. Building upon the success of the first edition, this c…

cs.CV2026

The First Challenge on Remote Sensing Infrared Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

Kai Liu, Haoyang Yue, Zeli Lin +65

This paper presents the NTIRE 2026 Remote Sensing Infrared Image Super-Resolution (x4) Challenge, one of the associated challenges of NTIRE 2026. The challenge aims to recover high…

eess.AS2026

Silent Speech Interfaces in the Era of Large Language Models: A Comprehensive Taxonomy and Systematic Review

Kele Xu, Yifan Wang, Ming Feng +5

Human-computer interaction has traditionally relied on the acoustic channel, a dependency that introduces systemic vulnerabilities to environmental noise, privacy constraints, and…