collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games

Zhenhao Chen, Yongqiang Chen, Chenxi Liu +7

Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundamentally relies on uncovering causal relati…

cs.CL2026

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Hongjian Zhou, Xinyu Zou, Jinge Wu +19

Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increa…

cs.CL2026

Pretraining Language Models on Historical Text

Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber +5

We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data qua…

cs.CL2026

The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models

Bohang Sun, Max Zhu, Francesco Caso +5

Diffusion large language models promise faster generation by refining many token positions in parallel, but this parallelism introduces a hidden control problem: which proposed tok…

cs.CL2026

TDGNet: Hallucination Detection in Diffusion Language Models via Temporal Dynamic Graphs

Arshia Hemmat, Philip Torr, Yongqiang Chen +1

Diffusion language models (D-LLMs) offer parallel denoising and bidirectional context, but hallucination detection for D-LLMs remains underexplored. Prior detectors developed for a…

cs.CL2025

TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models

Shenxu Chang, Junchi Yu, Weixing Wang +4

Diffusion large language models (D-LLMs) have recently emerged as a promising alternative to auto-regressive LLMs (AR-LLMs). However, the hallucination problem in D-LLMs remains un…