artificial intelligence

VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence

arXiv:2607.14510

summary

The paper introduces VLT, a multimodal foundation model that jointly processes industrial time‑series data, their frequency‑spectrum visualizations, and textual information to improve robustness and generalization in prognostics and health management tasks.

Abstract

Industrial time series serve as the foundation for Prognostics and Health Management (PHM) to ensure the reliability and safety of industrial equipment such as aero-engines. However, existing approaches are typically limited to single-modality modeling, which restricts their generalization in complex scenarios. Although recent advances in large language models (LLMs) provide new opportunities for multimodal learning, bridging continuous time-series signals and discrete textual semantics remains an open challenge. To this end, we propose VLT, a multimodal foundation model that jointly models time-series, frequency-spectrum visual representations, and textual knowledge. A key insight is to utilize the frequency spectrum as a visual bridge to connect continuous temporal signals with discrete semantics. Specifically, a Time-aware Mixture-of-Experts (Time-MoE) is designed to capture heterogeneous temporal dynamics, while a Frequency-Text Augmented Learner enables joint modeling of spectral and semantic features within a shared representation space. Furthermore, a time-centric gradient alignment mechanism is introduced to mitigate cross-modal optimization conflicts via gradient normalization and reliability-aware dynamic reweighting. Extensive experiments on multiple industrial datasets demonstrate that VLT outperforms state-of-the-art methods, achieving superior robustness and generalization under few-shot, noisy, and incomplete-modality settings.

18 pages, 13 figures, and 13 tables, including supplementary material. Haiteng Wang and Jingheng Yan contributed equally to this work

Topics & keywords

#multimodal learning#time series analysis#industrial intelligence#foundational models#few-shot learningTime-aware Mixture-of-ExpertsFrequency-Text Augmented Learnergradient alignmentfrequency spectrum visual bridgeprognostics and health management
VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence · wovepaper