4 papers
Scaffold Effects on GAIA: A Controlled Comparison
Jason Starace
Published agent capability scores conflate what a model can do with what its scaffold lets it do, and the magnitude of this elicitation gap is not well characterized under controll…
Ethical Implications of Training Deceptive AI
Jason Starace, Bert Baumgaertner, Terence Soule
Deceptive behavior in AI systems is no longer theoretical: large language models strategically mislead without producing false statements, maintain deceptive strategies through saf…
Intentional Deception as Controllable Capability in LLM Agents
Jason Starace, Terence Soule
As LLM-based agents increasingly operate in multi-agent systems, understanding adversarial manipulation becomes critical for defensive design. We present a systematic study of inte…
Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief Systems
Jason Starace, Terence Soule
How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach that infers agent properties from acti…