16 papers
TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments
Qingren Yao, Yaxuan Kong, Yuqi Nie +6
Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions. Existing ti…
Aionoscope: Debugging Latent-State Accessibility in Time-Series Representations
Alexander Chemeris, Ming Jin, Randall Balestriero
Time-series models are often evaluated by what they can forecast or classify, but those scores do not show whether their representations preserve the process state a user may want…
Large Models for Time Series and Spatio-Temporal Data: A Survey and Outlook
Ming Jin, Yaxuan Kong, Yuxuan Liang +13
Temporal data, including time series and spatio-temporal data, are pervasive in real-world applications. Generated in massive volumes by physical and virtual sensors, they record d…
It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks
Zhongzheng Qiao, Sheng Pan, Anni Wang +7
Time series foundation models (TSFMs) are revolutionizing the forecasting landscape from specific dataset modeling to generalizable task evaluation. However, we contend that existi…
TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
Yaxuan Kong, Qingren Yao, Yuqi Nie +7
Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural language and tools, it remains u…
SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents
Wentao Hu, Zhendong Chu, Yiming Zhang +6
Retrieval-augmented LLM agents increasingly rely on curated skill banks: collections of reusable textual principles that guide decision making on complex tasks. Existing approaches…