This chapter moves beyond risk registers to the governance structures, lifecycle processes, and evidence management practices required to make AI operations demonstrably accountable in a sovereign cloud context. It describes an AI governance operating model comprising an AI Governance Board, an AI Ethics Committee, a Model Risk Management function, and a Sovereign Compliance function, each with defined mandates and decision rights that integrate with existing enterprise governance bodies. The chapter covers model lifecycle governance from development through retirement, foundation model selection under sovereign constraints (including training data residency and inference location), policy-as-code enforcement through OPA/Rego, and bias detection across operational segments. Architects will find practical guidance on constructing a single structured evidence repository from which EU AI Act conformity documentation, DORA ICT risk management evidence, and NIS2 notification packages can be generated, reducing the manual overhead of multi-framework regulatory compliance.
The previous chapter established a taxonomy of risks associated with deploying AI agents in sovereign operational environments and described the control mechanisms—guardrails, autonomy tiers, human-in-the-loop checkpoints—through which those risks are mitigated at the point of execution. That chapter answered the question “what can go wrong, and how do we prevent it?” This chapter addresses a different and arguably more consequential question: “who decides, and how do we prove it?”
Risk identification and control implementation are necessary but insufficient conditions for trustworthy AI operations. A well-designed guardrail that prevents an agent from executing an unauthorised change is valuable, but it does not, by itself, answer the questions that a regulator, an auditor, or a board member will ask. Who approved the policy that defines what constitutes an unauthorised change? Who reviewed the model that the agent relies upon, and when was it last validated? What process governs the introduction of a new foundation model into the estate? How does the organisation ensure that its AI systems do not systematically disadvantage certain geographies, customer segments, or operational domains? Where is the evidence that these governance processes are not merely documented but actually followed?
These are governance questions, and they require governance structures—standing bodies with defined mandates, documented decision rights, repeatable lifecycle processes, and auditable evidence trails. This chapter describes how to construct an AI governance framework that is specific to the demands of sovereign cloud operations: one that respects jurisdictional boundaries, integrates with existing enterprise governance, leverages policy-as-code for enforcement, and produces the regulatory evidence that post-DORA, post-NIS2 supervisory regimes demand.

Every enterprise that has engaged seriously with AI risk management possesses, at minimum, a risk register—a catalogue of identified risks, their assessed likelihood and impact, the controls designated to mitigate them, and the residual risk accepted by the organisation. The risk taxonomy described in Chapter 21 provides a structured basis for such a register. Yet experience across regulated industries demonstrates that the risk register is, on its own, an inert artefact. It catalogues what is known about risk at a point in time but does nothing to ensure that the controls it describes are implemented, maintained, reviewed, and adapted as the environment changes [1].
The gap between a risk register and effective risk management is a governance gap. It manifests in several characteristic ways. First, risk registers age silently. A risk assessment performed during initial model deployment captures the risk profile as it existed at that moment; as the model is retrained, as the operational environment changes, as regulatory expectations evolve, the register drifts from reality unless a governance process mandates periodic reassessment. Second, risk registers are often disconnected from operational decision-making. The register may state that model drift is a medium-severity risk with monitoring as the designated control, but unless a governance body reviews drift metrics on a defined cadence and has the authority to suspend a model that exceeds drift thresholds, the control exists on paper but not in practice. Third, risk registers do not, by themselves, resolve contested decisions. When the AI engineering team believes a new agent capability is ready for production deployment and the compliance function disagrees, the risk register provides no mechanism for adjudication. A governance structure does.
The transition from risk registers to governance structures requires three foundational elements. Standing governance bodies with defined mandates, membership, and meeting cadences provide the organisational locus for AI-related decisions. Decision rights frameworks specify which decisions are made by which bodies, at which level, and under what constraints—preventing both governance paralysis (every decision escalated to the board) and governance bypass (consequential decisions made informally). Evidence management processes ensure that governance decisions, their rationale, and their outcomes are documented in a form that withstands regulatory scrutiny.
In a sovereign cloud context, each of these elements acquires an additional dimension. Governance bodies must account for jurisdictional variation: a model deployment decision that is straightforward in one sovereign zone may require additional review in another, because the regulatory regime differs or the data classification constraints are more stringent. Decision rights must be zone-aware: the authority to approve an agent for autonomous operation in a development environment does not automatically extend to a production environment in a regulated sovereign zone. Evidence management must satisfy multiple regulatory frameworks simultaneously—EU AI Act, DORA, NIS2, sector-specific regulations—each with different reporting formats, retention periods, and supervisory expectations.
The remainder of this chapter describes each component of the governance framework in detail, beginning with the operating model that defines who governs, and proceeding through the lifecycle processes, policy enforcement mechanisms, fairness requirements, and evidence management capabilities that make governance operational rather than aspirational.
An AI governance operating model defines the organisational structures, roles, responsibilities, and decision flows through which the enterprise exercises oversight of its AI capabilities. For an organisation operating a sovereign cloud estate with agentic operations, this model must integrate three dimensions: the AI-specific governance that addresses model risk, algorithmic fairness, and AI ethics; the IT governance that addresses change management, service management, and operational risk; and the sovereignty governance that addresses jurisdictional compliance, data residency, and cross-border data flows [2].
The AI Governance Board sits at the apex of the governance structure. Its mandate is to set the organisation’s AI strategy, approve the AI risk appetite, review material AI deployments, and ensure alignment between AI initiatives and the organisation’s regulatory obligations. The board should meet quarterly at minimum, with the authority to convene extraordinary sessions when material events—a significant model failure, a regulatory enforcement action, a change in the regulatory landscape—require urgent deliberation. Membership should include the Chief Information Officer or Chief Technology Officer, the Chief Risk Officer, the Chief Data Officer, the head of AI engineering, a representative from the legal function with AI regulatory expertise, and a representative from the business units most affected by AI-driven operations. In organisations where sovereign operations span multiple jurisdictions, a representative from each major jurisdictional compliance function should participate, at least on a rotational basis.
The AI Governance Board is a decision-making body, not merely an advisory one. Its decision rights include: approval of the organisation’s AI risk appetite statement; approval of new foundation model introductions into the estate; approval of agent autonomy tier elevations (the decision to permit an agent class to operate at a higher autonomy level, as defined in Chapter 21); review and acceptance of material model validation findings; and escalation decisions when model risk metrics breach defined thresholds. These decision rights should be formally documented in a board charter that is reviewed annually.

The AI Ethics Committee operates as a standing sub-committee of the AI Governance Board, with a more focused mandate: to evaluate AI deployments for ethical implications, review bias detection findings, assess the fairness of agent-driven operational decisions, and provide guidance on ethically ambiguous cases. The committee’s membership should extend beyond the technology and risk functions to include perspectives from human resources (particularly where AI decisions affect workforce management), customer-facing functions (where AI decisions affect service delivery equity), and, where practical, external advisors with expertise in AI ethics and algorithmic accountability. The committee’s output is advisory to the AI Governance Board, but its assessments should be required inputs to material deployment decisions.
The Model Risk Management (MRM) function is the technical governance capability responsible for independent model validation, ongoing model monitoring, and model risk reporting. In financial services, this function is well-established under SR 11-7 supervisory guidance [3]; in other sectors, it is increasingly required as AI deployments grow in materiality. The MRM function does not build or operate models—that is the responsibility of the AI engineering teams. Instead, it provides independent assurance that models perform as intended, that their limitations are understood and documented, that they are monitored for degradation, and that their risk profile remains within the boundaries set by the AI Governance Board. For sovereign operations, the MRM function must also validate that models operate within jurisdictional constraints: that inference occurs in the correct sovereign zone, that training data provenance is documented, and that model behaviour does not vary inappropriately across jurisdictional boundaries.
The Sovereign Compliance function ensures that AI governance decisions are consistent with the regulatory requirements of each sovereign zone in which the organisation operates. This function maintains the mapping between AI governance policies and jurisdictional regulations, identifies gaps or conflicts, and advises the AI Governance Board on zone-specific requirements. Where the EU AI Act classifies an AI system as high-risk, the Sovereign Compliance function ensures that the corresponding obligations—conformity assessment, registration in the EU database, ongoing monitoring—are met. Where DORA requires documented oversight of AI-related ICT services, the function ensures that the governance framework produces the evidence needed to demonstrate compliance.
The relationship between this AI governance structure and the existing IT governance structure is one of integration, not replacement. The Change Advisory Board (CAB) retains its authority over operational changes, including changes that involve AI components. The IT Risk Committee retains its oversight of operational risk, including risks arising from AI systems. What the AI governance structure adds is the specialised capability to evaluate risks and make decisions that require AI-specific expertise—model validation, algorithmic fairness assessment, foundation model selection—that the existing IT governance bodies are not equipped to perform. In practice, this means that a decision to deploy a new agent capability flows through both governance streams: the AI Governance Board approves the model and the agent’s autonomy tier, and the CAB approves the operational change that introduces the agent into the production environment.
A RACI matrix for AI agent deployment decisions clarifies these overlapping authorities. The AI engineering team is Responsible for developing and testing the agent capability. The MRM function is Accountable for independent validation. The AI Governance Board is Consulted on strategic alignment and risk appetite. The CAB is Informed of the deployment as an operational change. For sovereign zone deployments, the Sovereign Compliance function is Consulted on jurisdictional requirements. For deployments involving personal data processing, the Data Protection Officer is Consulted under GDPR and equivalent national legislation. This RACI structure prevents both the scenario where a consequential deployment proceeds without adequate review and the scenario where governance overhead renders deployment timelines impractical.
The AI governance operating model described in Section 22.2 provides the organisational structure for oversight. Model lifecycle governance provides the process framework: the defined stages, stage gates, approval workflows, and artefacts that ensure every model progresses from concept to retirement under consistent, documented governance.
The model lifecycle in a sovereign operations context comprises six stages: development, validation, deployment, monitoring, review, and retirement. Each stage has defined entry criteria, activities, exit criteria, and governance artefacts. The IBM watsonx.governance platform provides tooling to track models through this lifecycle, maintaining the lineage, metadata, and approval records that governance requires [4].
Development is the stage at which a model is built, trained, or fine-tuned for its intended operational purpose. The governance requirement at this stage is not to constrain the engineering process but to ensure that the model’s intended use, the data used for training, the performance objectives, and the known limitations are documented from the outset. The primary governance artefact produced during development is the model card—a structured document, following the template introduced by Mitchell et al. [5], that records the model’s purpose, training data provenance, performance metrics on representative evaluation sets, known failure modes, intended deployment context, and ethical considerations. For sovereign operations, the model card must additionally record the jurisdictional constraints that apply to the model’s deployment: which sovereign zones it is approved for, what data residency requirements its training data satisfies, and what inference location constraints apply.
The model card is not a one-time document. It is a living governance artefact that is updated at each lifecycle stage. watsonx.governance maintains the model card as a versioned record, linked to the model’s metadata in the AI model inventory, ensuring that any reviewer can trace the complete history of a model’s documentation from initial creation through every subsequent modification [4].
Validation is the stage gate at which independent assurance is provided that the model performs as intended and that its risk profile is acceptable. The MRM function conducts the validation, which typically includes: reproduction of performance metrics on held-out evaluation data; stress testing under adversarial or edge-case conditions; bias and fairness assessment (discussed in detail in Section 22.6); evaluation of the model’s behaviour at the boundaries of its intended operating envelope; and review of the model card for completeness and accuracy. The validation produces a validation report that documents findings, identifies any conditions or restrictions on deployment, and provides a recommendation to the AI Governance Board.
The AI Governance Board reviews the validation report and makes the deployment authorisation decision. This decision may be unconditional approval, conditional approval (deployment permitted subject to specified monitoring conditions or operational restrictions), or rejection (requiring the model to return to development for remediation). The decision, its rationale, and any conditions are recorded in the governance system and linked to the model’s lifecycle record in watsonx.governance.

Deployment is the operational transition from a validated model to a production-serving model. In a sovereign operations environment, deployment involves not only the technical provisioning of the model to the appropriate inference infrastructure but also the configuration of the agent that will use the model, the activation of monitoring instrumentation, and the registration of the deployment in the operational topology maintained by IBM Concert. The governance requirement at deployment is to verify that all pre-deployment conditions specified in the deployment authorisation have been met: that monitoring is active, that the agent’s autonomy tier is correctly configured, that the deployment is in the correct sovereign zone, and that the ITSM change record for the deployment has been approved by the CAB.
Monitoring is the continuous post-deployment stage during which the model’s performance, behaviour, and risk profile are tracked against the baselines established during validation. watsonx.governance provides automated monitoring for key metrics—prediction drift, feature drift, accuracy degradation, fairness metric deviation—and generates alerts when metrics breach configured thresholds [4]. Monitoring is not a passive activity; it is the mechanism through which the organisation detects model degradation before it manifests as operational impact. The monitoring stage produces periodic monitoring reports that are reviewed by the MRM function and, for material deviations, escalated to the AI Governance Board.
Review is the periodic reassessment of a deployed model’s continued suitability for its operational role. The review cadence is risk-proportionate: high-risk models (those operating at elevated autonomy tiers, those processing sensitive data, those deployed in regulated sovereign zones) are reviewed quarterly; lower-risk models are reviewed semi-annually or annually. The review considers monitoring data, incident history involving the model, changes in the operational environment that may affect the model’s performance, and changes in the regulatory landscape that may affect the model’s compliance status. The review produces a review finding that either confirms continued deployment, imposes additional conditions, or recommends retirement.
Retirement is the governed withdrawal of a model from operational service. Retirement may be triggered by a review finding, by the availability of a superior replacement model, or by a change in operational requirements. The governance requirement is to ensure that retirement is orderly: that dependent agents are reconfigured or decommissioned, that operational data produced by the model during its service life is retained in accordance with regulatory retention requirements, and that the retirement decision is documented in the model’s lifecycle record. A model that is simply switched off without governance is not retired; it is abandoned, and the distinction matters for audit and regulatory purposes.
The selection of foundation models for sovereign operations is not a purely technical decision. It is a governance decision with jurisdictional, regulatory, and strategic dimensions that must be evaluated through the governance structures described in Section 22.2. The proliferation of foundation models—proprietary models from commercial providers, open-weight models from research organisations and technology companies, and domain-specific models trained for particular industries—creates a selection problem that requires structured evaluation criteria rather than ad hoc technical benchmarking.
For a sovereign operations context, the evaluation criteria fall into five categories.
Data residency of training data is the first and often the most complex consideration. Foundation models are trained on large corpora, and the provenance of that training data determines whether the model can be deployed in sovereign zones with strict data origin requirements. A model trained on data that includes material from jurisdictions with incompatible data protection regimes may face deployment restrictions, particularly if the training data includes personal data that was processed without a legal basis recognised by the sovereign zone’s regulatory authority. The EU AI Act, in its requirements for high-risk AI systems, requires documentation of the data used for training, including information about data provenance and preparation [6]. For proprietary models where the training data composition is not fully disclosed by the provider, this documentation obligation may be impossible to satisfy—a governance risk that must be weighed in the selection process.
Inference location determines where the model executes when serving operational requests. In a sovereign operations architecture, inference must occur within the sovereign zone that is appropriate for the data being processed and the operational context being served. This means that the model must be deployable to infrastructure within the required jurisdiction—either on-premises, in a sovereign cloud region, or in a jurisdiction-appropriate hosted environment. Models that are available only as API services hosted in a single jurisdiction cannot satisfy this requirement for sovereign zones in other jurisdictions. IBM Granite models [7] are designed for flexible deployment: they are available both as hosted services on watsonx.ai and as downloadable model weights that can be deployed to on-premises or sovereign cloud infrastructure, providing the deployment flexibility that sovereign operations require.
Model hosting jurisdiction extends beyond inference location to encompass the legal jurisdiction that governs the model provider’s obligations. A model hosted by a provider subject to extraterritorial data access legislation—such as the US CLOUD Act—may present a sovereignty risk even if the inference infrastructure is physically located within the required sovereign zone, because the provider may be compelled to disclose data or model outputs to a foreign government. This risk is a governance consideration, not a technical one, and it must be evaluated by the Sovereign Compliance function in consultation with legal counsel.
Open-weight versus proprietary trade-offs represent a strategic governance decision. Open-weight models—those whose model weights are published and can be inspected, modified, and deployed independently—offer several governance advantages: the organisation can audit the model’s architecture, conduct independent security assessment, deploy the model to any infrastructure without vendor dependency, and retain full control over the model’s lifecycle. Proprietary models may offer superior performance on specific benchmarks but introduce governance dependencies: the organisation cannot inspect the model’s internals, cannot independently verify claims about training data provenance, cannot prevent the provider from modifying the model, and is subject to the provider’s terms of service which may change. The IBM Granite family occupies a position that addresses many of these concerns: Granite models are released under open licences with documented training data provenance, enabling organisations to deploy them within sovereign zones with full visibility into the model’s composition and lineage [7].
Operational fitness is the final evaluation criterion: the model’s performance on tasks representative of the operational workload it will serve. This criterion is evaluated through structured benchmarking against operational test cases—incident classification accuracy, remediation recommendation quality, anomaly detection sensitivity, natural language understanding of operational queries—rather than through generic benchmarks that may not reflect operational performance. The MRM function should maintain a library of operational test cases, drawn from the organisation’s incident history and operational experience, that is used consistently across model evaluations.
The foundation model selection process should produce a model selection decision record that documents the evaluation against each criterion, the governance body’s assessment, the conditions under which the selection is approved, and any restrictions on the model’s deployment. This record is maintained in the governance system and linked to the model’s lifecycle record, ensuring that the basis for the selection decision is available to auditors and regulators.

The governance structures and lifecycle processes described in the preceding sections define what governance decisions are made and by whom. Policy-as-code translates those decisions into machine-enforceable rules that are applied automatically at the point of deployment and operation, closing the gap between governance intent and operational reality [8].
The case for policy-as-code in AI governance is grounded in a simple observation: governance policies that rely solely on human compliance are governance policies that will be violated, not through malice but through the ordinary pressures of operational velocity. An engineer deploying a model under time pressure may overlook the requirement to verify that the deployment target is in the correct sovereign zone. A team releasing a new agent version may not check whether the agent’s autonomy tier has been approved for the target environment. A model that has been conditionally approved for deployment may be deployed without the conditions being met. Policy-as-code prevents these violations not by trusting human diligence but by encoding the governance rules into the deployment pipeline, where they are evaluated automatically and where non-compliant deployments are blocked before they reach production.
The Open Policy Agent (OPA) framework, with its Rego policy language, provides a mature, widely adopted foundation for policy-as-code in cloud-native environments [8]. OPA evaluates policy decisions by comparing a structured input (the deployment request, the model metadata, the target environment specification) against a set of policy rules and returning an allow or deny decision with an explanation. The policy rules are version-controlled in the same repositories as the infrastructure code, reviewed through the same pull request processes, and deployed through the same pipelines—ensuring that governance policy changes are traceable, auditable, and reversible.
For AI governance in sovereign operations, the following policy categories are essential.
Model deployment zone policies enforce the constraint that models may only be deployed to sovereign zones for which they have been approved. The policy rule evaluates the model’s approved zone list (maintained in the model’s governance metadata in watsonx.governance) against the target deployment zone specified in the deployment request. A model approved for Zone 2 (standard commercial operations) cannot be deployed to Zone 1 (regulated financial services) without a separate approval that updates the model’s governance metadata. The policy rule encodes this constraint directly:
package ai.governance.deployment
default allow = false
allow {
model := input.model_id
target_zone := input.target_zone
approved_zones := data.model_registry[model].approved_zones
target_zone == approved_zones[_]
data.model_registry[model].lifecycle_status == "deployed"
data.model_registry[model].validation_status == "approved"
}
Agent autonomy tier policies enforce the constraint that agents may not operate at an autonomy tier higher than that approved by the AI Governance Board for their agent class and target environment. An agent class approved for autonomy tier 2 (semi-autonomous with pre-execution human approval) in production cannot be elevated to tier 3 (autonomous with post-execution review) without a governance decision that updates the approved tier. The policy rule compares the requested autonomy tier against the approved maximum for the agent class and environment combination.
Mandatory human approval threshold policies enforce the requirement that certain categories of action—those above a defined risk threshold, those affecting regulated services, those involving cross-zone data movement—require explicit human approval regardless of the agent’s autonomy tier. These policies act as a floor beneath the autonomy tier system: even a tier 3 agent operating autonomously must obtain human approval for actions that exceed the defined thresholds.
Model validation currency policies enforce the requirement that a model’s validation must be current—that is, within the validation review period defined by the model’s risk classification—before the model may serve operational requests. A model whose most recent validation is older than the defined period (quarterly for high-risk models, annually for standard-risk models) is blocked from serving until a new validation is completed. This policy prevents the gradual erosion of validation currency that occurs when review cycles slip without consequence.
Training data provenance policies enforce the requirement that models deployed in sovereign zones with data origin restrictions have documented training data provenance that satisfies those restrictions. The policy evaluates the model’s training data provenance metadata against the sovereign zone’s data origin requirements, blocking deployment where provenance documentation is incomplete or non-compliant.
The policy rules are organised into a policy hierarchy that mirrors the governance structure. Organisation-wide policies, set by the AI Governance Board, form the base layer. Zone-specific policies, maintained by the Sovereign Compliance function, add jurisdictional constraints. Service-specific policies, maintained by service owners with AI Governance Board approval, add operational context. The policy evaluation engine evaluates all applicable layers for each decision, and a deployment is permitted only if all layers allow it. This layered structure ensures that zone-specific or service-specific policies can tighten but never relax organisation-wide policies—a principle that prevents local exceptions from undermining the governance framework.
Policy evolution is itself a governed process. Changes to policy rules are proposed through pull requests, reviewed by the governance function responsible for the policy layer being modified, tested against a suite of representative policy evaluation scenarios, and merged only after approval. The version history of every policy rule is maintained in the version control system, providing a complete audit trail of governance policy changes. watsonx.governance integrates with the policy-as-code pipeline to record which policy version was in effect at the time of each deployment decision, linking the deployment record to the specific policy rules that authorised it.

The discourse on AI bias has rightly focused on high-profile domains—hiring, lending, criminal justice—where the consequences of biased algorithmic decisions are direct and measurable. Operational AI in cloud environments might appear, at first glance, to be free from bias concerns: the systems being managed are technical infrastructure, and the decisions being made—whether to scale a service, how to prioritise an incident, which remediation to apply—seem to involve objective technical criteria rather than the subjective judgements that give rise to bias in human-facing applications.
This appearance is misleading. Operational AI systems make decisions that have differential impact across geographies, customer segments, and organisational units, and these decisions are shaped by training data and operational patterns that may encode systematic biases. Understanding how bias manifests in operational AI—and building the monitoring capability to detect it—is a governance obligation, not merely an ethical aspiration.
Consider three concrete scenarios. First, differential alert response by geography. An anomaly detection model trained predominantly on telemetry from well-instrumented European data centres may have lower sensitivity to anomalies in telemetry from African or Asian data centres where instrumentation patterns differ, network latency introduces different signal characteristics, or the volume of training data is smaller. The result is that incidents in some geographies are detected later, prioritised lower, or investigated less thoroughly—not because of a deliberate policy but because the model’s training data distribution creates a performance disparity. If the organisation serves customers in those geographies under service level agreements that promise equivalent service quality, the bias has both ethical and contractual consequences.
Second, biased incident prioritisation. An agent that prioritises incidents based on historical resolution patterns may learn to prioritise incidents affecting revenue-generating services over incidents affecting internal services or services used by smaller customer segments, even when the actual business impact is comparable. This bias arises from the historical data: revenue-generating service incidents attracted faster response in the past because of organisational attention patterns, and the model learns to replicate that pattern as if it were a rule. The effect is a systematic under-prioritisation of incidents whose impact falls on less commercially prominent stakeholders.
Third, remediation recommendation bias. A model that recommends remediation actions based on historical success patterns may systematically favour remediation approaches that were effective in one operational context (for example, public cloud environments with elastic scaling) while neglecting approaches that are more appropriate in another context (for example, on-premises sovereign environments with fixed capacity constraints). The result is that agents operating in sovereign zones receive less effective remediation recommendations—a form of operational bias that undermines the sovereign operations capability.
Detecting these biases requires a monitoring framework that evaluates model and agent performance across the dimensions that matter: geography, sovereign zone, service tier, customer segment, and operational context. The fairness metrics adapted for operational AI include equalised detection rates (the model’s anomaly detection sensitivity should be consistent across geographies and sovereign zones), equalised prioritisation (incidents of comparable business impact should receive comparable priority regardless of service category or customer segment), and equalised recommendation quality (remediation recommendations should achieve comparable success rates across operational contexts).
watsonx.governance provides a fairness monitoring capability that can be configured to track these metrics across defined segments [4]. The configuration requires the organisation to define the protected attributes (geography, sovereign zone, service tier) and the fairness metrics to be monitored, and to set thresholds beyond which a disparity triggers a governance review. When a threshold is breached, the monitoring system generates an alert to the MRM function, which investigates the cause and reports findings to the AI Ethics Committee. The committee assesses whether the disparity represents a genuine bias that requires remediation—model retraining, data augmentation, post-processing adjustment—or a legitimate variation that reflects genuine differences in the operational context.
The remediation of identified bias follows a structured process. Where bias arises from training data imbalance, the primary remedy is data augmentation—enriching the training set with representative data from underrepresented segments. Where bias arises from the model architecture or objective function, the remedy may involve retraining with a modified objective that incorporates fairness constraints. Where bias arises from the operational context rather than the model itself—for example, where sovereign zones with stricter network controls genuinely produce different telemetry characteristics—the remedy may be context-specific model tuning or the deployment of zone-specific model variants that are trained on locally representative data. In all cases, the remediation is documented, the model card is updated to reflect the identified bias and the corrective action taken, and the MRM function validates the effectiveness of the remediation before the model returns to production.
The EU AI Act provides additional regulatory impetus for bias monitoring. For AI systems classified as high-risk, Article 10 requires that training, validation, and testing data sets be subject to examination for possible biases, and Article 9 requires that risk management measures address risks of bias [6]. While most operational AI systems in cloud management may not be classified as high-risk under the current regulation, the governance framework should anticipate evolving classification criteria and build the monitoring capability proactively.

The governance framework described in the preceding sections generates a substantial body of evidence: governance board decisions, model cards, validation reports, deployment authorisations, monitoring reports, policy evaluation records, fairness assessments, and review findings. This evidence serves an internal governance purpose—demonstrating to the organisation’s own leadership that AI operations are governed effectively—but it also serves a regulatory purpose: providing the documented proof that supervisory authorities require under multiple overlapping regulatory frameworks.
The regulatory landscape for AI governance in European sovereign operations is defined by three principal instruments, each with distinct evidence requirements. The EU AI Act [6] requires providers and deployers of high-risk AI systems to maintain technical documentation, implement quality management systems, conduct conformity assessments, and register systems in the EU database. The documentation requirements are extensive: a description of the AI system, its intended purpose, the data governance measures applied to training data, the metrics used to measure accuracy, robustness, and cybersecurity, the human oversight measures, and the monitoring and logging capabilities. The Digital Operational Resilience Act (DORA) [9] requires financial entities to maintain a comprehensive ICT risk management framework that encompasses AI-related ICT services, to report major ICT-related incidents to competent authorities, and to conduct digital operational resilience testing. The Network and Information Security Directive (NIS2) [10] requires essential and important entities to implement risk management measures for network and information systems, to report significant incidents, and to ensure supply chain security—requirements that extend to AI components of the operational infrastructure.
Each regulatory instrument requires evidence, but the evidence formats, reporting timelines, and supervisory expectations differ. DORA incident reports must be submitted within prescribed timeframes (initial notification within four hours for major incidents in the financial sector). NIS2 early warnings must be submitted within 24 hours. EU AI Act conformity documentation must be maintained continuously and made available upon request. The governance framework must produce evidence that satisfies all applicable instruments without requiring separate evidence collection processes for each.
The approach is to design the governance framework’s evidence management as a single, structured evidence repository from which regulatory reports are generated through purpose-specific extraction and formatting. The evidence repository receives governance artefacts from across the framework: model cards from the development stage, validation reports from the MRM function, deployment authorisation records from the AI Governance Board, policy evaluation records from the OPA pipeline, monitoring reports from watsonx.governance, fairness assessment findings from the bias monitoring framework, and review findings from periodic reassessments. Each artefact is stored with structured metadata—artefact type, applicable model or agent, date, authoring body, applicable regulatory instruments—that enables targeted retrieval.
Automated evidence collection is essential for maintaining the evidence repository without imposing unsustainable manual overhead. The governance tooling—watsonx.governance for model lifecycle evidence, OPA for policy evaluation evidence, Concert for operational topology evidence, ServiceNow for ITSM process evidence—should be configured to generate and deposit governance artefacts automatically as governance processes execute. A model validation that is conducted through watsonx.governance automatically produces a validation record in the evidence repository. A policy evaluation that blocks a non-compliant deployment automatically produces a policy enforcement record. An agent action that triggers a human approval workflow automatically produces an approval record with the approver’s identity, the approval timestamp, and the action that was approved.
Audit packages are pre-assembled collections of evidence organised for a specific regulatory or audit purpose. An EU AI Act audit package for a specific AI system assembles the model card, training data documentation, validation report, deployment record, monitoring reports, and fairness assessments into a single, navigable collection that maps to the regulation’s documentation requirements. A DORA audit package for an ICT risk management review assembles the AI risk register entries, control evidence, incident records involving AI systems, and testing results into the format expected by the financial supervisor. The governance framework should include templates for the most commonly required audit packages, maintained by the Sovereign Compliance function and updated as regulatory expectations evolve.

Regulator engagement is the human process that accompanies the technical evidence management. The Sovereign Compliance function should maintain an active relationship with the relevant supervisory authorities, understanding their expectations, participating in industry consultations, and proactively sharing the governance framework’s approach when regulatory clarity is emerging. Regulators in the post-DORA, post-NIS2 landscape have demonstrated a preference for organisations that engage proactively and transparently over those that provide minimum compliance documentation only when required [9]. The governance framework’s evidence repository enables this proactive engagement by making it straightforward to assemble evidence packages that demonstrate the organisation’s governance maturity, even in advance of formal supervisory requests.
The evidence retention policy must account for the longest retention period required by any applicable regulation. DORA requires retention of ICT risk management documentation for at least five years. The EU AI Act requires that technical documentation be kept up to date and available for ten years after the AI system is placed on the market. NIS2 does not specify a retention period, but national transposition measures typically require retention for at least three years. The governance framework should apply a default retention period of ten years for all governance artefacts, with the ability to extend retention for specific artefacts where regulatory or legal requirements demand it.
Risk registers catalogue what is known about AI risk at a point in time but do not, on their own, ensure that controls are implemented, maintained, or adapted. The transition from risk management to effective governance requires standing governance bodies with defined mandates, documented decision rights, and auditable evidence trails.
The AI governance operating model comprises an AI Governance Board (strategic decisions and risk appetite), an AI Ethics Committee (fairness and ethical assessment), a Model Risk Management function (independent validation and monitoring), and a Sovereign Compliance function (jurisdictional regulatory alignment). These structures integrate with, rather than replace, existing IT governance bodies such as the Change Advisory Board.
Model lifecycle governance defines six stages—development, validation, deployment, monitoring, review, and retirement—each with stage gates, governance artefacts, and responsible roles. The model card, maintained as a versioned record in watsonx.governance, is the primary governance artefact that accompanies a model throughout its lifecycle.
Foundation model selection for sovereign operations must evaluate training data residency, inference location, hosting jurisdiction, open-weight versus proprietary trade-offs, and operational fitness. IBM Granite models address sovereign deployment requirements through open licences, documented training data provenance, and flexible deployment options including on-premises sovereign infrastructure.
Policy-as-code, implemented through OPA/Rego and integrated with the deployment pipeline, translates governance decisions into machine-enforceable rules—covering model deployment zones, agent autonomy tiers, mandatory human approval thresholds, validation currency, and training data provenance. Policy evolution is itself a governed, version-controlled process.
Bias in operational AI manifests through differential detection rates across geographies, biased incident prioritisation, and unequal remediation recommendation quality. Fairness monitoring across defined segments—geography, sovereign zone, service tier—is a governance obligation that watsonx.governance supports through configurable fairness metrics and threshold alerting.
Regulatory evidence management requires a single structured evidence repository from which EU AI Act conformity documentation, DORA ICT risk management evidence, and NIS2 incident notification packages are generated through purpose-specific extraction. Automated evidence collection from governance tooling and pre-assembled audit packages reduce the manual overhead of regulatory compliance.
This chapter has described the governance framework needed to operationalise AI risk management at enterprise scale in sovereign cloud environments. It has moved from the recognition that risk registers alone are insufficient, through the design of an AI governance operating model with defined bodies and decision rights, to the lifecycle processes that govern models from development through retirement. It has addressed the sovereign-specific challenge of foundation model selection, the enforcement of governance decisions through policy-as-code, the detection and remediation of bias in operational AI, and the evidence management architecture that satisfies multiple overlapping regulatory frameworks.
The governance framework described here provides the structures, processes, and evidence that regulators and auditors require. But governance is credible only to the extent that it is transparent—and transparency, in the context of AI-driven operational decisions, requires more than documented processes. It requires that the decisions AI systems make are explainable to the humans who must oversee them, and that the reasoning behind those decisions is auditable after the fact. Chapter 23 takes up this challenge directly, examining explainability and auditability as operational capabilities: how to make agent reasoning visible, how to construct audit trails that withstand regulatory scrutiny, and how to build the technical infrastructure that makes AI-driven operations not merely governed but genuinely understood by the humans who bear responsibility for them.
[1] National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, National Institute of Standards and Technology, Gaithersburg, MD, USA, Jan. 2023. [Online]. Available: https://doi.org/10.6028/NIST.AI.100-1
[2] International Organization for Standardization, “ISO/IEC 42001:2023 — Information technology — Artificial intelligence — Management system,” ISO/IEC, Geneva, Switzerland, 2023. [Online]. Available: https://www.iso.org/standard/81230.html
[3] Board of Governors of the Federal Reserve System, “Supervisory Guidance on Model Risk Management,” SR Letter 11-7, Board of Governors of the Federal Reserve System, Washington, DC, USA, Apr. 2011. [Online]. Available: https://www.federalreserve.gov/supervisionreg/srletters/sr1107.htm
[4] IBM, “IBM watsonx.governance: AI Governance and Model Risk Management,” IBM Documentation, IBM Corp., Armonk, NY, USA, 2024. [Online]. Available: https://www.ibm.com/docs/en/watsonx/saas?topic=governance
[5] M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru, “Model cards for model reporting,” in Proc. Conference on Fairness, Accountability, and Transparency (FAT)*, Atlanta, GA, USA, Jan. 2019, pp. 220–229. [Online]. Available: https://doi.org/10.1145/3287560.3287596
[6] European Parliament and Council of the European Union, “Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act),” Official Journal of the European Union, L Series, Jun. 2024. [Online]. Available: https://eur-lex.europa.eu/eli/reg/2024/1689
[7] IBM, “IBM Granite Models: Enterprise Foundation Models,” IBM Research, IBM Corp., Armonk, NY, USA, 2024. [Online]. Available: https://www.ibm.com/granite
[8] Open Policy Agent, “OPA: Open Policy Agent — Policy-based control for cloud native environments,” Cloud Native Computing Foundation, 2024. [Online]. Available: https://www.openpolicyagent.org/docs/latest/
[9] European Parliament and Council of the European Union, “Regulation (EU) 2022/2554 on digital operational resilience for the financial sector (DORA),” Official Journal of the European Union, L 333, pp. 1–79, Dec. 2022. [Online]. Available: https://eur-lex.europa.eu/eli/reg/2022/2554
[10] European Parliament and Council of the European Union, “Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2),” Official Journal of the European Union, L 333, pp. 80–152, Dec. 2022.
[11] AXELOS, ITIL 4: Practice Guide — Risk Management, AXELOS Ltd., London, UK, 2020.
[12] A. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes, “Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing,” in Proc. Conference on Fairness, Accountability, and Transparency (FAT)*, Barcelona, Spain, Jan. 2020, pp. 33–44.
Sovereign Cloud Operations: AI-Driven Management of Sovereign Estates
© 2026 by Alan Hamilton
is licensed under CC BY-SA 4.0