8 papers
VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation
Linwei Zhai, Han Ding, Mingzhi Lin +5
Vector Quantized Variational Autoencoders (VQ-VAEs) are fundamental to modern generative modeling, yet they often suffer from training instability and "codebook collapse" due to th…
WS-IMUBench: Can Weakly Supervised Methods from Audio, Image, and Video Be Adapted for IMU-based Temporal Action Localization?
Pei Li, Jiaxi Yin, Lei Ouyang +4
IMU-based Human Activity Recognition (HAR) has enabled a wide range of ubiquitous computing applications, yet its dominant clip classification paradigm cannot capture the rich temp…
What Matters in LLM-Based Feature Extractor for Recommender? A Systematic Analysis of Prompts, Models, and Adaptation
Kainan Shi, Peilin Zhou, Ge Wang +2
Using Large Language Models (LLMs) to generate semantic features has been demonstrated as a powerful paradigm for enhancing Sequential Recommender Systems (SRS). This typically inv…
We Can Hear You with mmWave Radar! An End-to-End Eavesdropping System
Dachao Han, Teng Huang, Han Ding +4
With the rise of voice-enabled technologies, loudspeaker playback has become widespread, posing increasing risks to speech privacy. Traditional eavesdropping methods often require…
Active Domain Adaptation for mmWave-based HAR via Renyi Entropy-based Uncertainty Estimation
Mingzhi Lin, Teng Huang, Han Ding +4
Human Activity Recognition (HAR) using mmWave radar provides a non-invasive alternative to traditional sensor-based methods but suffers from domain shift, where model performance d…
L3AC: Towards a Lightweight and Lossless Audio Codec
Linwei Zhai, Han Ding, Cui Zhao +4
Neural audio codecs have recently gained traction for their ability to compress high-fidelity audio and provide discrete tokens for generative modeling. However, leading approaches…