A Matter of Time: Revealing the Structure of Time in Vision-Language Models
arXiv:2510.19559 · doi:10.1145/3746027.3758163
Abstract
Large-scale vision-language models (VLMs) such as CLIP have gained popularity for their generalizable and expressive multimodal representations. By leveraging large-scale training data with diverse textual metadata, VLMs acquire open-vocabulary capabilities, solving tasks beyond their training scope. This paper investigates the temporal awareness of VLMs, assessing their ability to position visual content in time. We introduce TIME10k, a benchmark dataset of over 10,000 images with temporal ground truth, and evaluate the time-awareness of 37 VLMs by a novel methodology. Our investigation reveals that temporal information is structured along a low-dimensional, non-linear manifold in the VLM embedding space. Based on this insight, we propose methods to derive an explicit ``timeline'' representation from the embedding space. These representations model time and its chronological progression and thereby facilitate temporal reasoning tasks. Our timeline approaches achieve competitive to superior accuracy compared to a prompt-based baseline while being computationally efficient. All code and data are available at https://tekayanidham.github.io/timeline-page/.
References in corpus (11)
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Reproducible scaling laws for contrastive language-image learning
- Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Performance
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- Exploring CLIP for Assessing the Look and Feel of Images
- Do Language Models Understand Time?
- EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
- UniDiff: Advancing Vision-Language Models with Generative and Discriminative Learning
- Explaining Explainability: Recommendations for Effective Use of Concept Activation Vectors
- Blind Dates: Examining the Expression of Temporality in Historical Photographs
- Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express