Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
Rahul Baxi
Current language model evaluations measure what models know under ideal conditions but not how robustly they know it under realistic stress. Static benchmarks like MMLU and Truthfu…
cs.AI2026
The Comprehension-Gated Agent Economy: A Robustness-First Architecture for AI Economic Agency
Rahul Baxi
AI agents are increasingly granted economic agency (executing trades, managing budgets, negotiating contracts, and spawning sub-agents), yet current frameworks gate this agency on…