17 papers
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games
Zhenhao Chen, Yongqiang Chen, Chenxi Liu +7
Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundamentally relies on uncovering causal relati…
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Hongjian Zhou, Xinyu Zou, Jinge Wu +19
Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increa…
Pretraining Language Models on Historical Text
Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber +5
We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data qua…
Plan2Map: A Multimodal Benchmark for Document-Grounded Geospatial Boundary Reconstruction from Planning Records
Fabian Degen, Oishi Deb, Jindong Gu +4
Planning records define restrictions over geographic areas, but their source documents often provide only indirect spatial evidence rather than machine-readable boundaries. We intr…
The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer
Yupeng Chen, Junchi Yu, Aoxi Liu +3
Recent advances in end-to-end trained omni-models have substantially improved audio capabilities by strengthening text-audio modality alignment. However, whether such alignment ina…
-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing
Aoxi Liu, Yupeng Chen, James Oldfield +5
Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring for D-LLMs remains largely…