4 papers
Expected Free Energy as Belief-Dependent Utility for rho-POMDPs
Patrick Cooper, Alvaro Velasquez
An agent acting under partial observability must decide when to gather information and which observations are worth their cost. Standard POMDPs value information only through its e…
Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
Patrick Cooper, Alvaro Velasquez
Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stake in the outcome) and uncertainty suppres…
DeFAb: A Verifiable Benchmark for Defeasible Abduction in Foundation Models
Patrick Cooper, Alvaro Velasquez
A rule-based logic solver resolves every instance in our benchmark in under 50 microseconds with 100% accuracy; the best frontier language model reaches 65% at best and drops to 23…
Monotonicity as an Architectural Bias for Robust Language Models
Patrick Cooper, Alireza Nadali, Ashutosh Trivedi +1
Large language models (LLMs) are known to exhibit brittle behavior under adversarial prompts and jailbreak attacks, even after extensive alignment and fine-tuning. This fragility r…