3 papers
cs.CV2025
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models
Alexis Roger, Prateek Humane, Daniel Z. Kaplan +7
The proliferation of Vision-Language Models (VLMs) in the past several years calls for rigorous and comprehensive evaluation methods and benchmarks. This work analyzes existing VLM…
cs.AI2024
Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent
Karolis Jucys, George Adamopoulos, Mehrab Hamidi +7
Understanding the mechanisms behind decisions taken by large foundation models in sequential decision making tasks is critical to ensuring that such systems operate transparently a…
cs.LG2023
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
Kashif Rasul, Arjun Ashok, Andrew Robert Williams +15
Over the past years, foundation models have caused a paradigm shift in machine learning due to their unprecedented capabilities for zero-shot and few-shot generalization. However,…