3 papers
cs.AI2026
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
Malgorzata Gwiazda, Yifu Cai, Mononito Goswami +2
Large Language Models (LLMs) have shown promising performance in time series modeling tasks, but do they truly understand time series data? While multiple benchmarks have been prop…
cs.LG2026
Feynman: Knowledge-Infused Diagramming Agent for Scalable Visual Designs
Zixin Wen, Yifu Cai, Kyle Lee +5
Visual design is an essential application of state-of-the-art multi-modal AI systems. Improving these systems requires high-quality vision-language data at scale. Despite the abund…
cs.LG2025
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
Yifu Cai, Xinyu Li, Mononito Goswami +3
We introduce TimeSeriesGym, a scalable benchmarking framework for evaluating Artificial Intelligence (AI) agents on time series machine learning engineering challenges. Existing be…