9 papers
Lip-Siri: Contactless Open-Sentence Silent Speech with Wi-Fi Backscatter
Ye Tian, Haohua Du, Chao Gu +5
Silent speech interfaces (SSIs) enable silent interaction in noise-sensitive or privacy-sensitive settings. However, existing SSIs face practical deployment trade-offs among privac…
FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting
Yafei Lyu, Hao Zhou, Lu Zhang +2
Time series forecasting is central to data analysis and web technologies. The recent success of Large Language Models (LLMs) offers significant potential for this field, especially…
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
Lu Wang, Hao Chen, Siyu Wu +5
Multimodal Large Language Models (MLLMs) have been widely applied in speech and music. This tendency has led to a focus on audio tokenization for Large Models (LMs). Unlike semanti…
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
Haoran Zhou, Xingchen Song, Brendan Fahy +9
OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence…
Enabling Versatile Controls for Video Diffusion Models
Xu Zhang, Hao Zhou, Haoming Qin +5
Despite substantial progress in text-to-video generation, achieving precise and flexible control over fine-grained spatiotemporal attributes remains a significant unresolved challe…
Erase Diffusion: Empowering Object Removal Through Calibrating Diffusion Pathways
Yi Liu, Hao Zhou, Wenxiang Shang +2
Erase inpainting, or object removal, aims to precisely remove target objects within masked regions while preserving the overall consistency of the surrounding content. Despite diff…