6 papers
Notes to Self: Can LLMs Benefit from Experiential Abstractions?
Chang Liu, Xinyu Li, Artur Dubrawski
Humans distill experience into reusable abstractions, e.g., strategies and cautionary reminders, and apply them to gradually solve problems more effectively. We study whether Large…
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
Malgorzata Gwiazda, Yifu Cai, Mononito Goswami +2
Large Language Models (LLMs) have shown promising performance in time series modeling tasks, but do they truly understand time series data? While multiple benchmarks have been prop…
Investigating Compositional Reasoning in Time Series Foundation Models
Willa Potosnak, Cristian Challu, Mononito Goswami +4
Large pre-trained time series foundation models (TSFMs) have demonstrated promising zero-shot performance across a wide range of domains. However, a question remains: Do TSFMs succ…
Mitigating Persistent Client Dropout in Asynchronous Decentralized Federated Learning
Ignacy StÄpka, Nicholas Gisolfi, Kacper TrÄbacz +1
We consider the problem of persistent client dropout in asynchronous Decentralized Federated Learning (DFL). Asynchronicity and decentralization obfuscate information about model u…
Exploring Representations and Interventions in Time Series Foundation Models
MichaÅ WiliÅski, Mononito Goswami, Willa Potosnak +2
Time series foundation models (TSFMs) promise to be powerful tools for a wide range of applications. However, their internal representations and learned concepts are still not well…
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
Yifu Cai, Xinyu Li, Mononito Goswami +3
We introduce TimeSeriesGym, a scalable benchmarking framework for evaluating Artificial Intelligence (AI) agents on time series machine learning engineering challenges. Existing be…