From the 1 of 8 linked papers with an AI index.
8 papers
Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric
Oleg Solozobov
The paper introduces a vendor‑neutral metric that quantifies how reconstructable the evidence behind agent‑safety evaluation claims is, and provides tools (including Evidence Suffi…
DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency
Oleg Solozobov
Agent-runtime systems emit traces, ledgers, provenance graphs, policy logs, delegation tokens, cache events, and tool-firewall records, but those containers do not necessarily answ…
Property-Level Reconstructability of Agent Decisions: An Anchor-Level Pilot Across Vendor SDK Adapter Regimes
Oleg Solozobov
Agentic AI failures need post-hoc reconstruction: what the agent did, on whose authority, against which policy, and from what reasoning. Cross-regime feasibility remains unmeasured…
Decision Evidence Maturity Model for Agentic AI: A Property-Level Method Specification
Oleg Solozobov
Agentic AI systems produce decision evidence at scale through execution telemetry, but property-level reconstruction often fails when an external party asks a specific governance q…
Governed Auditable Decisioning Under Uncertainty: Synthesis and Agentic Extension
Oleg Solozobov
When automated decision systems fail, organizations frequently discover that formally compliant governance infrastructure cannot reconstruct what happened or why. This paper synthe…
Label-Free Detection of Governance Evidence Degradation in Risk Decision Systems
Oleg Solozobov
Risk decision systems in fraud detection and credit scoring operate under structural label absence: ground truth arrives weeks to months after decisions are made. During this blind…