papers

Publications (10)

cs.CV2025

Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context

Anatole Jacquin de Margerie, Alexis Roger, Irina Rish

Reproducibility remains a cornerstone of scientific progress, yet complex multimodal models often lack transparent implementation details and accessible training infrastructure. In…

cs.LG2025

Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models

Alexis Roger, Gwen Legate, Kashif Rasul +2

Tokenization and transfer learning are two critical components in building state of the art time series foundation models for forecasting. In this work, we systematically study the…

cs.CV2024

Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques

Rishika Bhagwatkar, Shravan Nayak, Reza Bayat +4

Vision-Language Models (VLMs) have witnessed a surge in both research and real-world applications. However, as they are becoming increasingly prevalent, ensuring their robustness a…

cs.CV2022

Aligning MAGMA by Few-Shot Learning and Finetuning

Jean-Charles Layoun, Alexis Roger, Irina Rish

The goal of vision-language modeling is to allow models to tie language understanding with visual inputs. The aim of this paper is to evaluate and align the Visual Language Model (…

cs.CL2025

Multilingual VLM Training: Adapting an English-Trained VLM to French

Jules Lahmi, Alexis Roger

Artificial intelligence has made great progress in recent years, particularly in the development of Vision--Language Models (VLMs) that understand both visual and textual data. How…

cs.CL2025

Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting

Roland Riachi, Kashif Rasul, Arjun Ashok +5

Recent works have demonstrated the effectiveness of adapting pre-trained language models (LMs) for forecasting time series in the low-data regime. We build upon these findings by a…

cs.AI2024

Towards ethical multimodal systems

Alexis Roger, Esma Aïmeur, Irina Rish

Generative AI systems (ChatGPT, DALL-E, etc) are expanding into multiple areas of our lives, from art Rombach et al. [2021] to mental health Rob Morris and Kareem Kouddous [2022];…

cs.CV2025

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models

Alexis Roger, Prateek Humane, Daniel Z. Kaplan +7

The proliferation of Vision-Language Models (VLMs) in the past several years calls for rigorous and comprehensive evaluation methods and benchmarks. This work analyzes existing VLM…

cs.CY2022

A review of modern surveillance techniques and their presence in our society

Alexis Roger

Technology is now omnipresent around us. Especially with the recent health crisis, many people started working remotely, bringing home an additional computer. Combining this with o…

cs.LG2026

LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series

Alexis Roger, Prateek Humane, Zhenghan Tai +4

Can language-pretrained transformers become effective time-series forecasters, and why? In this paper, we show that cross-modal transfer arises because language pretraining precond…