4 papers
The Forecast Critic: Leveraging Large Language Models for Poor Forecast Identification
Luke Bhan, Hanyu Zhang, Andrew Gordon Wilson +2
Monitoring forecasting systems is critical for customer satisfaction, profitability, and operational efficiency in large-scale retail businesses. We propose The Forecast Critic, a…
"Check My Work?": Measuring Sycophancy in a Simulated Educational Context
Chuck Arvin
This study examines how user-provided suggestions affect Large Language Models (LLMs) in a simulated educational context, where sycophancy poses significant risks. Testing five dif…
Identifying Legal Holdings with LLMs: A Systematic Study of Performance, Scale, and Memorization
Chuck Arvin
As large language models (LLMs) continue to advance in capabilities, it is essential to assess how they perform on established benchmarks. In this study, we present a suite of expe…
LLMForecaster: Improving Seasonal Event Forecasts with Unstructured Textual Data
Hanyu Zhang, Chuck Arvin, Dmitry Efimov +5
Modern time-series forecasting models often fail to make full use of rich unstructured information about the time series themselves. This lack of proper conditioning can lead to ob…