collaborators

5 papers

eess.AS2026

Noisy Environment Adaptation of Neural Speech Codec via Focal Mask and Noise Feature Separation

Shaokai Li, Weiping Tu, Yuhong Yang

Neural speech codec has attracted extensive attention for high-quality reconstruction at low-bitrate. However, real-world noise severely degrades its performance and hinders high-q…

cs.CV2026

GEM-TFL: Bridging Weak and Full Supervision for Forgery Localization through EM-Guided Decomposition and Temporal Refinement

Xiaodong Zhu, Yuanming Zheng, Suting Wang +4

Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments within videos or audio streams, providing interpretable evidence for multimedia forensics and se…

cs.CV2026

DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery Localization

Xiaodong Zhu, Suting Wang, Yuanming Zheng +5

Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments in video and audio, offering strong interpretability for security and forensics. While recent St…

cs.SD2025

FreeCodec: A disentangled neural speech codec with fewer tokens

Youqiang Zheng, Weiping Tu, Yueteng Kang +5

Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as spe…

cs.SD2025

Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model

Jizhen Li, Weiping Tu, Yuhong Yang +3

Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to subs…