5 papers
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Yadi Cao, Sicheng Lai, Jiahe Huang +12
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@…
CaTS-Bench: Can Language Models Describe Time Series?
Luca Zhou, Pratham Yashwante, Marshall Fisher +4
Time series captioning, the task of describing time series in natural language, requires numeric and temporal reasoning, trend interpretation, and contextual understanding. Existin…
Can LLMs Understand Time Series Anomalies?
Zihao Zhou, Rose Yu
Large Language Models (LLMs) have gained popularity in time series forecasting, but their potential for anomaly detection remains largely unexplored. Our study investigates whether…
Multi-Modal Forecaster: Jointly Predicting Time Series and Textual Data
Kai Kim, Howard Tsai, Rajat Sen +5
Current forecasting approaches are largely unimodal and ignore the rich textual data that often accompany the time series due to lack of well-curated multimodal benchmark dataset.…
Back to Bayesics: Uncovering Human Mobility Distributions and Anomalies with an Integrated Statistical and Neural Framework
Minxuan Duan, Yinlong Qian, Lingyi Zhao +4
Existing methods for anomaly detection often fall short due to their inability to handle the complexity, heterogeneity, and high dimensionality inherent in real-world mobility data…