3 papers
cs.LG2026
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
Hui Dai, Ryan Teehan, Parsa Torabian +1
Probabilistic forecasting estimates the likelihood of uncertain future events. To improve LLM forecasting, existing methods typically learn from binary outcomes to output verbalize…
cs.CL2025
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Hui Dai, Ryan Teehan, Mengye Ren
Many existing evaluation benchmarks for Large Language Models (LLMs) quickly become outdated due to the emergence of new models and training data. These benchmarks also fall short…
cs.CL2024
DENIAHL: In-Context Features Influence LLM Needle-In-A-Haystack Abilities
Hui Dai, Dan Pechi, Xinyi Yang +2
The Needle-in-a-haystack (NIAH) test is a general task used to assess language models' (LMs') abilities to recall particular information from long input context. This framework how…