collaborators

6 papers

cs.AI2026

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

Katherine Elkins, Jon Chun

Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framing. We audit 16 models across 14…

cs.AI2026

The Paradox of Robustness: Decoupling Rule-Based Logic from Affective Noise in High-Stakes Decision-Making

Jon Chun, Katherine Elkins

While Large Language Models (LLMs) are widely documented to be sensitive to minor prompt perturbations and prone to sycophantic alignment, their robustness in consequential, rule-b…

cs.CL2026

Syntactic Framing Fragility: An Audit of Robustness in LLM Ethical Decisions

Katherine Elkins, Jon Chun

Large language models exhibit systematic negation sensitivity, yet no operational framework exists to measure this vulnerability at deployment scale, especially in high-stakes deci…

cs.GT2026

Auditing Game-Theoretic Measures of Strategic Reasoning in LLMs

Mateo Pechon-Elkins, Jon Chun

Game-theoretic benchmarks can separate distinct forms of strategic behavior in LLMs that aggregate theory of mind scores do not distinguish, but only if those benchmarks measure wh…

cs.CL2026

CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models

Jon Chun, Hannah Sussman, Adrian Mangine +13

Pragmatic reasoning, inferring intended meaning beyond literal semantics, underpins everyday communication yet remains difficult for large language models. We present the Contextua…

cs.AI2026

AgenticSimLaw: A Juvenile Courtroom Multi-Agent Debate Simulation for Explainable High-Stakes Tabular Decision Making

Jon Chun, Kathrine Elkins, Yong Suk Lee

We introduce AgenticSimLaw, a role-structured, multi-agent debate framework that provides transparent and controllable test-time reasoning for high-stakes tabular decision-making t…