Spherity AI Safety Series — Paper 2 of 3

Risk-Bounded Runtime Assurance for Dynamic Multi-Agent Systems under Partial Observation

AI Safety: Verifiable Identity and Evidence for Bounded Cross-Organisation Autonomy

Author
Affiliation
Dr. Carsten Stöcker — Spherity GmbH
Published
Updated
Research cut-off
· Research and implementation evidence reviewed to this date
This version
https://spherity.github.io/spherity-research/risk-bounded-runtime-assurance-multi-agent-systems.html
Latest version
https://spherity.github.io/spherity-research/risk-bounded-runtime-assurance-multi-agent-systems.html
Browse this paper

In brief

What does this research establish?

The paper defines a protected runtime supervisor that admits an AI-agent action only when current mandate and an action-bound assurance lease both hold. Under stated calibration and model-coverage assumptions, a persistent episode account bounds first-harm probability even as agents join, retire or are replaced and effects remain pending.

Key takeaways

  • Current authority and runtime assurance are complementary: neither can substitute for the other at the execution boundary.
  • Episode risk, shared dependencies and pending effects must persist across worker replacement; resetting a worker must not reset responsibility.
  • Assurance leases bind evidence, action, time, model state and continuation allowance instead of granting open-ended approval.
  • The theorem is conditional on calibration and model coverage and is not a measured field-safety guarantee.

Download Paper 2

Download the full paper

The runtime supervisor, assurance leases, first-harm theorem, synthetic scenarios, state exploration and omission counterexamples.

Start with the executive brief

Download the executive brief

Read the leadership case and how all three technical papers form one AI Safety control chain.

First page of Risk-Bounded Runtime Assurance for Dynamic Multi-Agent Systems under Partial Observation.
AI Safety Series — Paper 2. Risk-bounded runtime assurance for dynamic multi-agent systems under partial observation. Published under CC BY 4.0.

Abstract

This paper develops a protected runtime supervisor for dynamic multi-agent systems operating with incomplete information. The supervisor admits an external action only when the actor holds a current mandate and an action-bound assurance lease remains valid for the exact proposal. It retains evidence, shared dependencies, pending effects and a persistent episode-risk account as workers join, retire or are replaced. Under explicit calibration and model-coverage assumptions, the paper derives a bound on first-harm probability and tests the controller with synthetic scenarios, state exploration and omission counterexamples.

Keywords: runtime assurance; AI safety; multi-agent systems; partial observation; assurance lease; episode risk; verifiable identity; bounded autonomy

The runtime assurance problem

Authorization can be valid while operating evidence becomes stale. A model can be capable while the environment moves outside its assessed envelope. A worker can be replaced while its earlier actions remain pending. Static approval therefore cannot govern a dynamic cross-company episode.

The proposed supervisor treats every protected external action as a new admission decision. It evaluates the current mandate, the evidence view, the action, remaining risk allowance and unresolved outcomes. This is a stateful control problem, not a one-time model certification.

Current authority plus action-bound assurance

A protected gateway admits an AI-agent action only when both current mandate and action-bound assurance are valid.
Figure 1. Current mandate and action-bound assurance meet at a protected execution gateway. Identity or a proof alone does not establish the hidden operating state; both eligibility inputs must be current for the exact action. Source: Spherity GmbH.

A valid mandate answers who may act and within which scope. An assurance lease answers whether a specified action remains justified under current evidence, time, model state and continuation allowance. The protected gateway requires both. This prevents an identity credential from being mistaken for safety evidence, and prevents an assurance result from granting legal or organizational authority.

Partial observation and hidden operating conditions

A partially observed supervisor maintains a belief over hidden operating conditions and the probability of undetected prior harm.
Figure 2. The real operating condition is partly hidden, while the supervisor acts on an evidence-based belief and a bound for undetected prior harm. A quiet or signed report can update belief but cannot reveal hidden state by itself. Source: Spherity GmbH.

The supervisor cannot observe every relevant condition directly. It maintains a belief over a finite model bank and separately accounts for the possibility that harm has already occurred without detection. Signed reports improve provenance but do not make the hidden state observable. Silence is evidence only under an explicitly calibrated observation model.

This distinction matters for remote industrial systems, independent model providers and cross-company data releases. The assurance layer must state what a signal can establish, how quickly evidence becomes stale and what conservative response applies when coverage is uncertain.

Persistent episode risk, pending effects and worker churn

The episode account permanently debits the continuation bound associated with admitted actions. It also retains dependencies and pending effects until verified closure. A worker may retire, crash or be replaced, but its contribution to episode exposure remains. That property prevents identity rotation or process churn from resetting the safety budget.

Suitable operational controls include action-specific leases, bounded time-to-live, idempotent operation IDs, verified effect closure, independent incident signals and a kill switch whose authority, scope and latency are tested rather than assumed.

Deployment responsibilities and use cases

Provider-managed, independently operated and hybrid cloud-edge AI deployments assign assurance, enforcement and response duties differently.
Figure 10. Provider-managed, independently operated and hybrid cloud-edge deployments place responsibility in different locations. Each protected effect still needs an eligibility gate, persistent episode accounting, oversight and a local response path. Source: Spherity GmbH.

The architecture assigns responsibility instead of centralizing all authority. Each resource owner maintains a protected local decision while the episode account preserves shared constraints and evidence.

Evaluation evidence and limits

The paper evaluates 432 synthetic scenarios comprising 120,060 terminal histories. The maximum reported first-harm rate is 3.73% against a 5% test threshold. A controller abstraction reaches 24,448 states, and twelve omission variants produce counterexamples.

These are design-validation results, not field calibration. They depend on the modeled state space, observation semantics, calibration quality, independence assumptions and complete mediation of relevant effects. Unmodeled dynamics, correlated evidence failures and physical hazards require separate treatment.

Selected references

  1. NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0).
  2. NIST, Zero Trust Architecture, SP 800-207.
  3. IETF, RFC 9334: Remote ATtestation procedureS Architecture.
  4. SPIFFE, SPIFFE Concepts.
  5. W3C, Verifiable Credentials Data Model v2.0.
  6. W3C, Decentralized Identifiers v1.0.

How to cite Paper 2

Stöcker, Carsten (2026). Risk-Bounded Runtime Assurance for Dynamic Multi-Agent Systems under Partial Observation: AI Safety—Verifiable Identity and Evidence for Bounded Cross-Organisation Autonomy. Spherity GmbH. https://spherity.github.io/spherity-research/risk-bounded-runtime-assurance-multi-agent-systems.html. Licensed CC BY 4.0.

Creative Commons Attribution 4.0 International

Open research

License and citation

This research page and the linked Paper 2 PDF is licensed by its named author under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Reuse must credit every named author, link to this canonical version and the license, and indicate whether changes were made.

How to cite this work

Dr. Carsten Stöcker (2026-09-28). “Risk-Bounded Runtime Assurance for Dynamic Multi-Agent Systems under Partial Observation: AI Safety: Verifiable Identity and Evidence for Bounded Cross-Organisation Autonomy.” Spherity GmbH. https://spherity.github.io/spherity-research/risk-bounded-runtime-assurance-multi-agent-systems.html. Licensed CC BY 4.0.

Direct answers

Questions this research answers

What is risk-bounded runtime assurance for AI agents?

It is a protected admission process that evaluates current authority, action-specific evidence, a partially observed operating state and remaining episode risk before permitting an external effect, then updates the persistent account from outcomes.

Why must episode risk survive worker replacement?

Replacing an agent process does not remove earlier exposure, shared dependencies or effects that are still pending. Resetting the risk account with the worker would let a system spend the same continuation allowance repeatedly.

Does the paper prove that a deployed multi-agent system is safe?

No. The first-harm bound is conditional on calibration, model-bank coverage, protected enforcement and other stated assumptions. The synthetic evaluation tests the controller logic but does not replace field measurement or general alignment evidence.