collaborators

9 papers

cs.CV2026

ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer

Jinyi Hu, Shengding Hu, Yuxuan Song +6

Autoregressive and diffusion models have achieved remarkable progress in language models and visual generation, respectively. We present ACDiT, a novel Autoregressive blockwise Con…

cs.HC2026

Lip-Siri: Contactless Open-Sentence Silent Speech with Wi-Fi Backscatter

Ye Tian, Haohua Du, Chao Gu +5

Silent speech interfaces (SSIs) enable silent interaction in noise-sensitive or privacy-sensitive settings. However, existing SSIs face practical deployment trade-offs among privac…

cs.SD2025

MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment

Hao Zhou, Xiaobao Guo, Yuzhe Zhu +1

Propelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual…

cs.LG2025

FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting

Yafei Lyu, Hao Zhou, Lu Zhang +2

Time series forecasting is central to data analysis and web technologies. The recent success of Large Language Models (LLMs) offers significant potential for this field, especially…

cs.CL2025

Assessing the Macro and Micro Effects of Random Seeds on Fine-Tuning Large Language Models

Nghia Bui, Guergana Savova, Lijing Wang

The impact of random seeds in fine-tuning large language models (LLMs) has been largely overlooked despite its potential influence on model performance.In this study, we systematic…

cs.SD2025

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation

Lu Wang, Hao Chen, Siyu Wu +5

Multimodal Large Language Models (MLLMs) have been widely applied in speech and music. This tendency has led to a focus on audio tokenization for Large Models (LMs). Unlike semanti…