9 papers
ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer
Jinyi Hu, Shengding Hu, Yuxuan Song +6
Autoregressive and diffusion models have achieved remarkable progress in language models and visual generation, respectively. We present ACDiT, a novel Autoregressive blockwise Con…
Lip-Siri: Contactless Open-Sentence Silent Speech with Wi-Fi Backscatter
Ye Tian, Haohua Du, Chao Gu +5
Silent speech interfaces (SSIs) enable silent interaction in noise-sensitive or privacy-sensitive settings. However, existing SSIs face practical deployment trade-offs among privac…
MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment
Hao Zhou, Xiaobao Guo, Yuzhe Zhu +1
Propelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual…
FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting
Yafei Lyu, Hao Zhou, Lu Zhang +2
Time series forecasting is central to data analysis and web technologies. The recent success of Large Language Models (LLMs) offers significant potential for this field, especially…
Assessing the Macro and Micro Effects of Random Seeds on Fine-Tuning Large Language Models
Nghia Bui, Guergana Savova, Lijing Wang
The impact of random seeds in fine-tuning large language models (LLMs) has been largely overlooked despite its potential influence on model performance.In this study, we systematic…
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
Lu Wang, Hao Chen, Siyu Wu +5
Multimodal Large Language Models (MLLMs) have been widely applied in speech and music. This tendency has led to a focus on audio tokenization for Large Models (LMs). Unlike semanti…