8 papers
Do Language Models Update their Forecasts with New Information?
Zhangdie Yuan, Zifeng Ding, Andreas Vlachos
Prior work has largely treated forecasting as a static task, failing to consider how forecasts and the confidence in them should evolve as new evidence emerges. To address this gap…
Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
Zifeng Ding, Shenyang Huang, Zeyu Cao +11
Forecasting future links is a central task in temporal graph (TG) reasoning, requiring models to leverage historical interactions to predict upcoming ones. Traditional neural appro…
TCP: a Benchmark for Temporal Constraint-Based Planning
Zifeng Ding, Sikuan Yan, Zhangdie Yuan +3
Temporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of comp…
Toward Reliable Clinical Coding with Language Models: Verification and Lightweight Adaptation
Zhangdie Yuan, Han-Chin Shing, Mitch Strong +1
Accurate clinical coding is essential for healthcare documentation, billing, and decision-making. While prior work shows that off-the-shelf LLMs struggle with this task, evaluation…
FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark
Zhangdie Yuan, Zifeng Ding, Andreas Vlachos
Forecasting is an important task in many domains, such as technology and economics. However existing forecasting benchmarks largely lack comprehensive confidence assessment, focus…
PRobELM: Plausibility Ranking Evaluation for Language Models
Zhangdie Yuan, Eric Chamoun, Rami Aly +2
This paper introduces PRobELM (Plausibility Ranking Evaluation for Language Models), a benchmark designed to assess language models' ability to discern more plausible from less pla…