collaborators

6 papers

cs.AI2026

Expected Free Energy as Belief-Dependent Utility for rho-POMDPs

Patrick Cooper, Alvaro Velasquez

An agent acting under partial observability must decide when to gather information and which observations are worth their cost. Standard POMDPs value information only through its e…

cs.LG2026

Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization

Patrick Cooper, Alvaro Velasquez

Discovering causal relationships requires controlled experiments, but experimentalists face a sequential decision problem: each intervention reveals information that should inform…

cs.MA2026

See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents

Siyi Chen, Xiaoyan Zhang, Meng Wu +7

Multi-agent systems communicate mostly through text, paying a lossy and expensive decode and re-encode cost. KV-cache communication is a promising alternative, yet most prior work…

cs.LG2026

Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization

Yancheng Huang, Changsheng Wang, Chongyu Fan +7

Foundation models, such as large language models (LLMs), are powerful but often require customization before deployment to satisfy practical constraints such as safety, privacy, an…

cs.LG2026

HIPO: Instruction Hierarchy via Constrained Reinforcement Learning

Keru Chen, Jun Luo, Sen Lin +4

Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO…

cs.CL2026

Monotonicity as an Architectural Bias for Robust Language Models

Patrick Cooper, Alireza Nadali, Ashutosh Trivedi +1

Large language models (LLMs) are known to exhibit brittle behavior under adversarial prompts and jailbreak attacks, even after extensive alignment and fine-tuning. This fragility r…