4 papers
WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning
Yizhou Chi, Eric Chamoun, Zifeng Ding +1
Forecasting real-world events requires language-model agents to reason under uncertainty from incomplete, time-bounded information. Yet evaluating whether agents genuinely forecast…
SciPaths: Forecasting Pathways to Scientific Discovery
Eric Chamoun, Yizhou Chi, Yulong Chen +4
Scientific progress depends on sequences of enabling contributions, yet existing AI4Science benchmarks largely focus on citation prediction, literature retrieval, or idea generatio…
Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP Artefacts
Eric Chamoun, Nedjma Ousidhoum, Michael Schlichtkrull +1
Clarifying the research framing of NLP artefacts (e.g., models, datasets, etc.) is crucial to aligning research with practical applications. Recent studies manually analyzed NLP re…
PRobELM: Plausibility Ranking Evaluation for Language Models
Zhangdie Yuan, Eric Chamoun, Rami Aly +2
This paper introduces PRobELM (Plausibility Ranking Evaluation for Language Models), a benchmark designed to assess language models' ability to discern more plausible from less pla…