activity
20242026
collaborators

11 papers

cs.AI2026

Demonstrating Generalization Failures via Mixtures of Conditional Policies

Jou Barzdukas, Jack Peck, Julian Schulz +3

Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training and deployment environments. This exposes…

cs.CY2026

Open Technical Problems in Open-Weight AI Model Risk Management

Stephen Casper, Kyle O'Brien, Shayne Longpre +19

Frontier AI models with openly available weights are steadily becoming more powerful and widely adopted. However, compared to proprietary models, open-weight models pose different…

cs.CL2026

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

Ajmal M., Abin Roy, Afthab Salam Kanniyan +4

Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation…

cs.CL2026

TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models

Jinho Choo, JunSeung Lee, Jimyeong Kim +3

Large language models (LLMs) demonstrate strong multilingual capabilities, yet often fail to consistently generate responses in the intended language, exhibiting a phenomenon known…

cs.LG2025

Depth-Wise Activation Steering for Honest Language Models

Gracjan Góral, Marysia Winkels, Steven Basart

Large language models sometimes assert falsehoods despite internally representing the correct answer, failures of honesty rather than accuracy, which undermines auditability and sa…

cs.LG2025

Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity

Austin Meek, Eitan Sprejer, Iván Arcuschin +2

Chain-of-thought (CoT) outputs let us read a model's step-by-step reasoning. Since any long, serial reasoning process must pass through this textual trace, the quality of the CoT i…