works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
most citedQA-LIGN: Aligning LLMs through Constitutionally Decomposed QA

2 citations · 2 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

Xiao Ye, Jacob Dineen, Evan Zhu +3

The paper presents Hindcast, a framework that evaluates large language model forecasters by replaying resolved prediction markets using a frozen Reddit snapshot taken before each m…

cs.CL2026

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

Jacob Dineen, Aswin RRV, Zhikun Xu +1

Co-evolutionary self-play, where one language model generates problems and another solves them, promises curriculum learning without human supervision. The promise breaks down earl…

cs.CL2026

RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems

Adarsh Srinivasan, Jacob Dineen, Muhammad Umar Afzal +3

Large language models in healthcare often produce emotionally flat or opaque responses, failing to provide the transparent reasoning required for clinical trust. We present RECAP (…

cs.CL20252 cited

QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA

Jacob Dineen, Aswin RRV, Qin Liu +8

Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the tra…

cs.CL2025

Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications

Xiao Ye, Jacob Dineen, Zhaonan Li +11

Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a…

cs.CL2025

ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation

Qin Liu, Jacob Dineen, Yuxi Huang +4

Benchmarks are central to measuring the capabilities of large language models and guiding model development, yet widespread data leakage from pretraining corpora undermines their v…