Spherity AI Safety Series — Paper 2 of 3
Risk-Bounded Runtime Assurance for Dynamic Multi-Agent Systems under Partial Observation
AI Safety: Verifiable Identity and Evidence for Bounded Cross-Organisation Autonomy
Browse this paper
In brief
What does this research establish?
The paper defines a protected runtime supervisor that admits an AI-agent action only when current mandate and an action-bound assurance lease both hold. Under stated calibration and model-coverage assumptions, a persistent episode account bounds first-harm probability even as agents join, retire or are replaced and effects remain pending.
Key takeaways
- Current authority and runtime assurance are complementary: neither can substitute for the other at the execution boundary.
- Episode risk, shared dependencies and pending effects must persist across worker replacement; resetting a worker must not reset responsibility.
- Assurance leases bind evidence, action, time, model state and continuation allowance instead of granting open-ended approval.
- The theorem is conditional on calibration and model coverage and is not a measured field-safety guarantee.
Download Paper 2
Download the full paperThe runtime supervisor, assurance leases, first-harm theorem, synthetic scenarios, state exploration and omission counterexamples.
Start with the executive brief
Download the executive briefRead the leadership case and how all three technical papers form one AI Safety control chain.
Abstract
This paper develops a protected runtime supervisor for dynamic multi-agent systems operating with incomplete information. The supervisor admits an external action only when the actor holds a current mandate and an action-bound assurance lease remains valid for the exact proposal. It retains evidence, shared dependencies, pending effects and a persistent episode-risk account as workers join, retire or are replaced. Under explicit calibration and model-coverage assumptions, the paper derives a bound on first-harm probability and tests the controller with synthetic scenarios, state exploration and omission counterexamples.
Keywords: runtime assurance; AI safety; multi-agent systems; partial observation; assurance lease; episode risk; verifiable identity; bounded autonomy
The runtime assurance problem
Authorization can be valid while operating evidence becomes stale. A model can be capable while the environment moves outside its assessed envelope. A worker can be replaced while its earlier actions remain pending. Static approval therefore cannot govern a dynamic cross-company episode.
The proposed supervisor treats every protected external action as a new admission decision. It evaluates the current mandate, the evidence view, the action, remaining risk allowance and unresolved outcomes. This is a stateful control problem, not a one-time model certification.
Current authority plus action-bound assurance
A valid mandate answers who may act and within which scope. An assurance lease answers whether a specified action remains justified under current evidence, time, model state and continuation allowance. The protected gateway requires both. This prevents an identity credential from being mistaken for safety evidence, and prevents an assurance result from granting legal or organizational authority.
Partial observation and hidden operating conditions
The supervisor cannot observe every relevant condition directly. It maintains a belief over a finite model bank and separately accounts for the possibility that harm has already occurred without detection. Signed reports improve provenance but do not make the hidden state observable. Silence is evidence only under an explicitly calibrated observation model.
This distinction matters for remote industrial systems, independent model providers and cross-company data releases. The assurance layer must state what a signal can establish, how quickly evidence becomes stale and what conservative response applies when coverage is uncertain.
Persistent episode risk, pending effects and worker churn
The episode account permanently debits the continuation bound associated with admitted actions. It also retains dependencies and pending effects until verified closure. A worker may retire, crash or be replaced, but its contribution to episode exposure remains. That property prevents identity rotation or process churn from resetting the safety budget.
Suitable operational controls include action-specific leases, bounded time-to-live, idempotent operation IDs, verified effect closure, independent incident signals and a kill switch whose authority, scope and latency are tested rather than assumed.
Deployment responsibilities and use cases
- Provider-managed systems place more assurance and resource control within the platform, but customers still need clear acceptance criteria and outcome evidence.
- Independently operated systems require portable identity, evidence and incident semantics because no single provider owns the full control loop.
- Hybrid cloud-edge systems must coordinate remote models and local protected effects; the edge cannot outsource immediate response to an unavailable cloud service.
- Cross-company data release can treat disclosure as an external effect, requiring mandate, purpose, evidence and episode constraints before release.
The architecture assigns responsibility instead of centralizing all authority. Each resource owner maintains a protected local decision while the episode account preserves shared constraints and evidence.
Evaluation evidence and limits
The paper evaluates 432 synthetic scenarios comprising 120,060 terminal histories. The maximum reported first-harm rate is 3.73% against a 5% test threshold. A controller abstraction reaches 24,448 states, and twelve omission variants produce counterexamples.
These are design-validation results, not field calibration. They depend on the modeled state space, observation semantics, calibration quality, independence assumptions and complete mediation of relevant effects. Unmodeled dynamics, correlated evidence failures and physical hazards require separate treatment.
Selected references
- NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0).
- NIST, Zero Trust Architecture, SP 800-207.
- IETF, RFC 9334: Remote ATtestation procedureS Architecture.
- SPIFFE, SPIFFE Concepts.
- W3C, Verifiable Credentials Data Model v2.0.
- W3C, Decentralized Identifiers v1.0.
How to cite Paper 2
Stöcker, Carsten (2026). Risk-Bounded Runtime Assurance for Dynamic Multi-Agent Systems under Partial Observation: AI Safety—Verifiable Identity and Evidence for Bounded Cross-Organisation Autonomy. Spherity GmbH. https://spherity.github.io/spherity-research/risk-bounded-runtime-assurance-multi-agent-systems.html. Licensed CC BY 4.0.
This research page and the linked Paper 2 PDF is licensed by its named author under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Reuse must credit every named author, link to this canonical version and the license, and indicate whether changes were made.
How to cite this work
Dr. Carsten Stöcker (2026-09-28). “Risk-Bounded Runtime Assurance for Dynamic Multi-Agent Systems under Partial Observation: AI Safety: Verifiable Identity and Evidence for Bounded Cross-Organisation Autonomy.” Spherity GmbH. https://spherity.github.io/spherity-research/risk-bounded-runtime-assurance-multi-agent-systems.html. Licensed CC BY 4.0.
Direct answers
Questions this research answers
What is risk-bounded runtime assurance for AI agents?
It is a protected admission process that evaluates current authority, action-specific evidence, a partially observed operating state and remaining episode risk before permitting an external effect, then updates the persistent account from outcomes.
Why must episode risk survive worker replacement?
Replacing an agent process does not remove earlier exposure, shared dependencies or effects that are still pending. Resetting the risk account with the worker would let a system spend the same continuation allowance repeatedly.
Does the paper prove that a deployed multi-agent system is safe?
No. The first-harm bound is conditional on calibration, model-bank coverage, protected enforcement and other stated assumptions. The synthetic evaluation tests the controller logic but does not replace field measurement or general alignment evidence.