Publications (10)
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
Anatole Jacquin de Margerie, Alexis Roger, Irina Rish
Reproducibility remains a cornerstone of scientific progress, yet complex multimodal models often lack transparent implementation details and accessible training infrastructure. In…
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
Alexis Roger, Gwen Legate, Kashif Rasul +2
Tokenization and transfer learning are two critical components in building state of the art time series foundation models for forecasting. In this work, we systematically study the…
Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques
Rishika Bhagwatkar, Shravan Nayak, Reza Bayat +4
Vision-Language Models (VLMs) have witnessed a surge in both research and real-world applications. However, as they are becoming increasingly prevalent, ensuring their robustness a…
Aligning MAGMA by Few-Shot Learning and Finetuning
Jean-Charles Layoun, Alexis Roger, Irina Rish
The goal of vision-language modeling is to allow models to tie language understanding with visual inputs. The aim of this paper is to evaluate and align the Visual Language Model (…
Multilingual VLM Training: Adapting an English-Trained VLM to French
Jules Lahmi, Alexis Roger
Artificial intelligence has made great progress in recent years, particularly in the development of Vision--Language Models (VLMs) that understand both visual and textual data. How…
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
Roland Riachi, Kashif Rasul, Arjun Ashok +5
Recent works have demonstrated the effectiveness of adapting pre-trained language models (LMs) for forecasting time series in the low-data regime. We build upon these findings by a…
Towards ethical multimodal systems
Alexis Roger, Esma Aïmeur, Irina Rish
Generative AI systems (ChatGPT, DALL-E, etc) are expanding into multiple areas of our lives, from art Rombach et al. [2021] to mental health Rob Morris and Kareem Kouddous [2022];…
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models
Alexis Roger, Prateek Humane, Daniel Z. Kaplan +7
The proliferation of Vision-Language Models (VLMs) in the past several years calls for rigorous and comprehensive evaluation methods and benchmarks. This work analyzes existing VLM…
A review of modern surveillance techniques and their presence in our society
Alexis Roger
Technology is now omnipresent around us. Especially with the recent health crisis, many people started working remotely, bringing home an additional computer. Combining this with o…
LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series
Alexis Roger, Prateek Humane, Zhenghan Tai +4
Can language-pretrained transformers become effective time-series forecasters, and why? In this paper, we show that cross-modal transfer arises because language pretraining precond…