works on

From the 1 of 17 linked papers with an AI index.

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2026

Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

Xiao Ye, Jacob Dineen, Evan Zhu +3

The paper presents Hindcast, a framework that evaluates large language model forecasters by replaying resolved prediction markets using a frozen Reddit snapshot taken before each m…

cs.CL2026

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

Jacob Dineen, Aswin RRV, Zhikun Xu +1

Co-evolutionary self-play, where one language model generates problems and another solves them, promises curriculum learning without human supervision. The promise breaks down earl…

cs.CL2026

Robust Asynchronous Planning via Auto-Formalization

Jiayi Zhang, Jianing Yin, Ben Zhou +1

LLMs can plan by either generating action sequences directly as a Planner or translating tasks into domain specific language for an external solver as a Formalizer. While most real…

cs.CL2026

RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems

Adarsh Srinivasan, Jacob Dineen, Muhammad Umar Afzal +3

Large language models in healthcare often produce emotionally flat or opaque responses, failing to provide the transparent reasoning required for clinical trust. We present RECAP (…

cs.CL2026

Reliable Use of Lemmas via Eligibility Reasoning and SectionAware Reinforcement Learning

Zhikun Xu, Xiaodong Yu, Ben Zhou +6

Recent large language models (LLMs) perform strongly on mathematical benchmarks yet often misapply lemmas, importing conclusions without validating assumptions. We formalize lemma$…

cs.CL2025

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

Matthew W. Kenaston, Umair Ayub, Mihir Parmar +14

Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology…