activity
20242026
collaborators

8 papers

cs.CL2026

Do Language Models Update their Forecasts with New Information?

Zhangdie Yuan, Zifeng Ding, Andreas Vlachos

Prior work has largely treated forecasting as a static task, failing to consider how forecasts and the confidence in them should evolve as new evidence emerges. To address this gap…

cs.AI2025

Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning

Zifeng Ding, Shenyang Huang, Zeyu Cao +11

Forecasting future links is a central task in temporal graph (TG) reasoning, requiring models to leverage historical interactions to predict upcoming ones. Traditional neural appro…

cs.AI2025

TCP: a Benchmark for Temporal Constraint-Based Planning

Zifeng Ding, Sikuan Yan, Zhangdie Yuan +3

Temporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of comp…

cs.CL2025

Toward Reliable Clinical Coding with Language Models: Verification and Lightweight Adaptation

Zhangdie Yuan, Han-Chin Shing, Mitch Strong +1

Accurate clinical coding is essential for healthcare documentation, billing, and decision-making. While prior work shows that off-the-shelf LLMs struggle with this task, evaluation…

cs.LG2025

FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark

Zhangdie Yuan, Zifeng Ding, Andreas Vlachos

Forecasting is an important task in many domains, such as technology and economics. However existing forecasting benchmarks largely lack comprehensive confidence assessment, focus…

cs.CL2025

PRobELM: Plausibility Ranking Evaluation for Language Models

Zhangdie Yuan, Eric Chamoun, Rami Aly +2

This paper introduces PRobELM (Plausibility Ranking Evaluation for Language Models), a benchmark designed to assess language models' ability to discern more plausible from less pla…