This chapter examines how agentic operations must integrate with IT Service Management disciplines to remain governable, traceable, and trusted at enterprise scale, arguing that agents must operate through ITSM processes rather than around them. It maps the four ITIL 4 practices most affected by agent participation—incident management, change enablement, problem management, and service request management—to concrete patterns in which Concert, Orchestrate, and multi-agent pipelines interact with ServiceNow, Jira, and GitHub. The chapter demonstrates how agent-assisted workflows produce consistently richer operational records, with every investigative and remediation action logged as an attributed entry on the ITSM record in real time. Architects will find guidance on tiered change management (standard, normal, and emergency), agent-driven root cause correlation across Concert’s knowledge graph, and the construction of a unified operational record that satisfies DORA, NIS2, and ISO 27001 evidence requirements as a natural by-product of correct operation.
The previous chapter established the architectural substrate for multi-agent orchestration: the role taxonomy, the communication mechanisms, the shared-context model, and the guardrails that allow agents to operate safely inside regulated sovereign zones. What that chapter deliberately did not address was the institutional layer that sits above the technical substrate—the managed processes, the accountability structures, and the documented records through which an organisation demonstrates control to regulators, auditors, customers and its own leadership. This chapter addresses that layer. It examines how an agentic operations capability, however sophisticated its internal architecture, must integrate with IT Service Management disciplines if it is to be governable, traceable, and trusted at enterprise scale.
There is a version of the automation story that presents ITSM as a relic—a body of process designed for a slower, more manual era, inevitably to be swept aside as agents take over operations. In this version, the change advisory board gives way to a continuous deployment pipeline; the incident ticket gives way to an AI-generated event record; the problem management process gives way to automatic root cause analysis. The story is seductive, and parts of it contain genuine insight. But taken as a prescription, it creates serious problems for any organisation that operates under regulatory obligation, and it misunderstands what ITSM actually does.
ITSM, in its contemporary form, is not primarily about slowing things down. It is about providing organisational accountability, cross-team coordination, and documented evidence of control. These three functions are not diminished by automation; they are made more complex by it, and therefore more in need of explicit design.
Consider the accountability dimension first. When a human engineer restarts a service, there is, at minimum, a person who can be asked why. When an agent restarts a service as part of an automated remediation workflow, the question of accountability becomes both more important and harder to answer without deliberate architecture. Who authorised the workflow? Under what policy did the agent act? Was the action reviewed before execution? The answers to these questions do not emerge automatically from the fact that automation was used. They must be constructed deliberately through the ITSM records—incident tickets, change requests, problem records—that constitute the organisation’s operational memory. Without those records, automated operations are not more accountable than manual ones. They are less accountable, because the chain of human decisions that would normally anchor accountability has been replaced by opaque automation.
The regulatory dimension makes this concrete. The Digital Operational Resilience Act (DORA) [1] requires financial entities to maintain a comprehensive incident register and to report major incidents to supervisors within prescribed timeframes. Article 17 of the regulation is specific: the incident management process must include procedures for recording, classifying and managing ICT-related incidents, with defined criteria for escalation and notification. NIS2 [2] imposes a comparable obligation on essential and important entities across eleven critical sectors, requiring documented incident notifications to national Computer Security Incident Response Teams within 24 hours of awareness. Neither regulation permits the organisation to substitute an AI-generated event log for a properly structured incident record. The record must be created, managed, and retained in a form that supervisors can inspect and auditors can rely upon. Nor does either regulation offer an exemption for organisations that have chosen to automate their operations. If anything, the introduction of AI agents into operational processes heightens the expectation of documented oversight: DORA’s Article 28 requires documented contractual arrangements for ICT third-party services, including AI service providers, and NIS2’s Article 20 imposes personal liability on management bodies for the adequacy of cybersecurity risk management measures.
ISO/IEC 27001:2022 [3] adds the information security management dimension. Clause A.5.37, which addresses documented operating procedures, and Clause A.8.32, which covers change management, both apply to automated operations just as they apply to manual ones. The standard does not regard “the automation does it” as a sufficient answer to the question of whether changes to information processing systems are being managed in a controlled and auditable way. It requires that changes be assessed, approved, tracked, tested, and documented, with roles and responsibilities clearly defined.
The risk of bypassing ITSM in an agentic environment is therefore not merely theoretical. Organisations that allow agents to take operational actions outside the structure of ITSM processes create accountability gaps that regulators interpret as evidence of inadequate control frameworks. In a post-DORA, post-NIS2 landscape, such gaps carry material consequences: supervisory intervention, mandatory remediation programmes, and, in cases of significant incidents that were not properly recorded or notified, potential financial penalties and management liability.
Cross-team coordination is the third reason ITSM remains essential. An agent handling an incident in the payment processing domain may need to coordinate with network engineering, security operations, platform engineering, and application development teams simultaneously. Without ITSM tooling as the shared coordination surface, each of those teams works from its own view of the situation. The ServiceNow incident record, or the equivalent in whichever ITSM platform the organisation uses, provides a single reference point: what is the current status, who is doing what, what has been tried, and who needs to be notified. Agents that operate outside this coordination surface—posting updates to Slack or generating internal logs that only they can read—create a fragmented picture that human operators cannot reconstruct reliably under pressure.
The relationship between agentic operations and ITSM is therefore not one of replacement but of integration. Agents can handle many of the mechanical tasks involved in ITSM—creating records, enriching them with telemetry, escalating through workflows, closing tickets with documented evidence—faster and more consistently than humans. But the records themselves, and the processes they support, remain the irreplaceable institutional artefact. The architecture must ensure that agent actions flow through ITSM rather than around it.
ITIL 4 [4] organises its guidance around thirty-four management practices, of which a subset are directly and substantially affected by the introduction of agentic operations. Four practices deserve particular attention: incident management, change enablement, problem management, and service request management. For each, the introduction of agents changes not only how work is performed but what the ITSM tooling must be capable of supporting.
Incident management in ITIL 4 is defined as the practice of minimising the negative impact of incidents by restoring normal service operation as quickly as possible [4]. The practice assumes a lifecycle: an incident is detected, recorded, classified, prioritised, investigated, resolved, and closed, with appropriate communication at each stage and a post-incident review for significant events. In a traditional operations model, the early stages of this lifecycle are heavily human-dependent: someone notices something wrong, creates a ticket, and begins investigation. Agent participation transforms the early stages. Agents can detect anomalies, correlate them with topology data, create the initial incident record, classify and prioritise based on service impact, and begin diagnostic actions—all before a human operator has been paged. This increases the speed and consistency of the early lifecycle, but it also changes what the ITSM tooling must handle. Records may now be created at machine speed, in large volumes, with structured machine-generated content. The incident priority field may now carry an agent-calculated severity score rather than a human estimate. The activity log within the ticket must capture agent actions with the same attribution granularity as human actions. And the escalation rules must account for situations where an agent has taken action that needs human review, not merely situations where a human has escalated from a lower tier.
Change enablement in ITIL 4 covers the practice of maximising the number of successful changes by ensuring that risks are properly assessed, that all changes are authorised, and that the change schedule is managed [4]. The practice distinguishes between standard changes (pre-approved, repeatable, low-risk), normal changes (requiring assessment and authorisation), and emergency changes (requiring fast-track authorisation when normal processes would take too long). Agents affect all three categories. For standard changes, an agent can execute the change autonomously once the pre-approval exists; the ITSM tooling must record the execution against the change model without requiring per-instance human approval. For normal changes, an agent can assemble and present the risk assessment that human reviewers need, enriching the change record with telemetry evidence, historical performance data, and topology impact analysis. For emergency changes, an agent can mobilise context faster than any human team, assembling the evidence required for a fast-track approval decision in minutes rather than hours. In each case, the ITSM tooling must accommodate machine-generated content in change records and support workflows that are partially automated and partially human-governed, sometimes within the same change record.
Problem management in ITIL 4 is the practice of reducing the likelihood and impact of incidents by identifying actual and potential causes of incidents and managing workarounds and known errors [4]. It operates at two speeds: reactive problem management, which responds to incidents that have already occurred; and proactive problem management, which identifies risks before they manifest. Agents contribute to both. In reactive mode, agents can correlate incident histories, identify common contributing factors, and generate structured problem records that human problem managers review and act upon. In proactive mode, agents can analyse telemetry trends and topology changes to identify configurations or behaviours that have preceded incidents in the past, flagging potential problems before they become incidents. The ITSM tooling must support problem records that include machine-generated analytical content—correlation matrices, statistical summaries of incident clustering, agent-identified potential root causes—alongside the human judgements that close or validate the analysis.
Service request management in ITIL 4 covers the handling of service requests from users: requests for information, access, resources, or standard service actions [4]. In a traditional model, service requests are handled manually by service desk staff or first-line support. Agents can automate the handling of many common request types entirely: access provisioning follows a defined workflow, certificate renewal is a scripted operation, environment creation is an infrastructure-as-code execution. Where automation cannot handle the full request, agents can handle the initial triage, data collection, and preliminary steps, leaving human operators with a context-enriched request that requires only a decision rather than a full investigation. The ITSM tooling must support service catalogue entries that are backed by agent-executed workflows, and it must track service level agreement compliance for requests handled at machine speed, where the volume and velocity are orders of magnitude higher than human capacity.
The common thread across all four practices is that agent participation changes the velocity, volume, and nature of the content flowing through ITSM processes, without changing the fundamental purpose of those processes. The implication for ITSM tooling is that platforms must be designed to consume and produce machine-readable, structured content at scale, to support mixed human-agent workflows within a single record, and to maintain the attribution and audit quality required by regulated enterprises. ServiceNow, as the market-leading ITSM platform in enterprise contexts, has invested substantially in these capabilities through its Now Automation suite and its integrations with AI operations platforms. The integration patterns described in this chapter assume ServiceNow as the primary ITSM surface, but the architectural principles apply to any platform with comparable API and workflow capabilities.
The incident management lifecycle in an agent-integrated environment can be described through seven stages. Each stage involves both agent activity and ITSM record-keeping, and the two are not separable: agent activity without ITSM records is ungoverned automation, and ITSM records without agent enrichment are incomplete artefacts that fail to capture the operational reality.
Detection is the point at which an anomaly or failure first becomes known to the operations control plane. In the architecture described in earlier chapters, detection is primarily the responsibility of IBM Concert [5], which continuously monitors the topology model and the signal streams flowing from Instana and other observability components. Concert’s correlation engine identifies patterns that deviate from established baselines—latency thresholds crossed, error rates spiking, dependency chains degraded—and generates a situation record when those patterns meet the configured significance criteria. The situation record is not yet an ITSM incident; it is Concert’s internal representation of a developing operational issue.
The transition from Concert situation to ITSM incident is the first integration point. A watsonx Orchestrate [6] agent, acting on a Concert event notification, calls the ServiceNow API to create an incident record, populating it with the situation identifier, the affected services identified by Concert’s topology model, the initial severity classification derived from Concert’s impact score, and a summary of the telemetry signals that triggered the situation. Crucially, the Concert situation identifier is written into the incident record as a cross-reference field, establishing the link between the ITSM record and the AI platform’s internal representation of the event. This link is the foundation of the unified operational record that subsequent sections will discuss.

Triage follows immediately after detection. An executor agent queries Concert’s knowledge graph to retrieve the full context of the affected services: their dependencies, their current health posture, the change history for the affected components in the preceding 72 hours, and the history of similar situations from Concert’s operational memory. This context enrichment takes, typically, less than two minutes for a well-configured Concert deployment against a known service topology. The resulting context is appended to the ServiceNow incident record as a structured enrichment note, and the incident priority is updated based on Concert’s calculated business impact. If the affected service is tier-one—a regulated financial transaction service, a critical citizen-facing application—the agent also triggers the notification workflow that pages the on-call team, supplying them with the pre-assembled context rather than a bare alert.
The triage agent also evaluates whether the situation matches a known error pattern recorded in the ServiceNow known error database. If a match is found, the agent appends the known error reference to the incident record and, where the known error includes an approved workaround, begins executing that workaround immediately. This is the point at which the investment in problem management, discussed in Section 19.5, pays operational dividends: known errors, properly documented and linked to approved workarounds, allow agents to begin remediation before a human has even read the incident notification.
Investigation is the stage at which agents pursue root cause hypotheses in parallel. A planner agent, using Concert’s impact topology, decomposes the investigation into sub-tasks: examine the database connection pool behaviour, check the network latency between the application tier and the storage layer, compare the current certificate validity status for the affected service’s mutual TLS configuration, and review the recent IaC deployment history for the affected infrastructure. Each sub-task is dispatched to an executor agent that has the credentials and tool access appropriate to that specific investigation domain. Executor agents report findings back to the planner via the shared working memory structure described in Chapter 18. The planner synthesises these findings into a structured hypothesis.
Every investigative action taken by an executor agent is logged as an activity entry on the ServiceNow incident record, attributed to the agent’s identity with a timestamp and a reference to the sub-task in the Concert workflow. A human operator opening the incident record at any point during the investigation sees a real-time activity feed that reads as coherently as if a single engineer were documenting their work, because the structured attribution logging enforces consistent formatting across all agent-generated entries.
Remediation follows the formation of a credible hypothesis. For low-risk remediations—service restarts, cache flushes, configuration rollbacks within a pre-approved range—a reviewer agent evaluates the proposed action against Open Policy Agent rules and, if the policy permits, the executor agent carries out the action. The action is recorded as a work note on the ServiceNow incident record, and if it involves a configuration change, a change request is automatically created in ServiceNow and linked to the incident. This link between the incident record and the change record is architecturally important: it allows auditors to trace, from the incident that motivated a change, through the change record that authorised it, to the specific configuration artefact that was modified. The chain is complete and machine-navigable without requiring anyone to assemble it after the fact.
For higher-risk remediations—failover between availability zones, modification of load balancer weights above a defined threshold, changes to security group rules—the reviewer agent withholds approval and creates a human approval task on the ServiceNow incident record. The task is routed to the appropriate approver based on the service’s ownership data in the Concert CMDB, and it includes the agent’s assembled evidence: the hypothesis, the proposed action, the expected outcome, the blast radius estimate, and the rollback procedure. The approver is not starting from scratch; they are reviewing a pre-assembled package and making a binary decision with the confidence that the agent has done the investigative work correctly. This is the human-in-the-loop checkpoint described in Chapter 18, expressed through an ITSM workflow rather than an API call.
Communication is managed by a dedicated communication agent that monitors the incident record and generates status updates for the relevant audiences at configured intervals. Internal updates are posted to the incident record and to the designated Slack or Teams channel. External updates, where the affected service has a public status page or a contractual SLA notification obligation, are generated from templates and reviewed by the on-call manager before publication. For regulated entities, NIS2’s 24-hour early warning obligation is tracked as a compliance timer on the incident record: if the incident meets the significance criteria and the timer is about to expire, the communication agent escalates to the incident manager with a pre-drafted notification for supervisory submission.
Closure is the stage most frequently handled with insufficient rigour in manual operations, because teams under pressure to resolve the next incident do not invest in documenting the last one adequately. Agent-assisted closure addresses this directly. Once the service metrics confirm restoration to normal operating parameters, a closure agent queries the full incident record—all activity notes, all change references, all agent-generated findings—and generates a structured incident summary that includes the timeline, the root cause hypothesis, the actions taken, the contributing change references, and the recommended follow-up actions for problem management. This summary is appended to the incident record before closure, creating the post-incident documentation that DORA’s Article 17 requires without depending on an engineer to write it under time pressure. The Concert situation is also updated and closed, maintaining synchronisation between the ITSM system and the AI platform’s operational record.

Change management in an agentic environment must balance two competing imperatives. On one side is the operational need for speed: agents that can detect, diagnose and propose a remediation in minutes should not be delayed by change governance processes designed around weekly CAB meetings. On the other side is the regulatory obligation for controlled change: DORA’s requirements for managed ICT risk, ISO 27001’s change management controls, and the organisational need to maintain an auditable record of what changed, when, and on whose authority. The resolution of this tension lies not in eliminating change governance but in designing it for the velocity at which agents operate.
ITIL 4’s three-tier classification of changes provides the structural framework. Standard changes are pre-approved: the risk has been assessed in advance, the procedure is documented, and no per-instance authorisation is required. In an agentic environment, standard changes are the category most amenable to full automation. An agent executing a standard change—rotating an API key within a permitted key management system, scaling a deployment within pre-approved bounds, restarting a service that has a documented restart procedure—creates a change record in ServiceNow, populates it with the executed procedure, the change model reference, the before-and-after state, and the validation evidence, and closes it. No human interaction is required during execution. Human oversight is present in the design of the change model: the CAB approves the standard change procedure once, and the agent’s execution is validated against that approved procedure at runtime by the policy engine. The change record is the evidence that this validation occurred.
The governance challenge with standard changes in an agentic environment is the risk of procedural drift: agents that execute changes that are substantively different from the approved procedure but that no individual human reviews closely enough to notice. The control for this is the combination of structured change models and post-execution sampling. Change models are defined in ServiceNow as structured templates with mandatory fields—target resource type, permitted parameter ranges, required validation steps—and the agent’s execution is constrained to the parameters defined in the model by the OPA guardrail layer. Post-execution, a fraction of completed standard changes are flagged for human review, with the agent-generated before-and-after state compared against the expected outcome specified in the change model. Persistent discrepancies trigger a review of the change model itself.
Normal changes require risk assessment and authorisation before execution. In a traditional model, the change initiator prepares a change record and a risk assessment, submits it to the CAB, and awaits approval. In an agentic model, the risk assessment is generated by Concert’s change risk scoring capability [5], which evaluates the proposed change against the topology model, the deployment history of the affected component, the current operational health of the dependent services, and the historical correlation between similar changes and subsequent incidents. The resulting risk score, with its supporting evidence, is written into the change record by an agent, providing the CAB with a structured, evidence-backed assessment rather than a human-authored narrative.
The CAB’s role in this model does not disappear; it shifts. Rather than evaluating unstructured risk narratives and making holistic judgements about whether sufficient care has been taken, CAB members review structured risk evidence, ask targeted questions about specific factors that the agent’s assessment has surfaced, and apply their judgement to the aspects of the change that are genuinely uncertain or contextually sensitive—factors that the agent’s model-based assessment cannot fully capture, such as business timing considerations, contractual commitments, or political sensitivities in a multi-tenant shared platform. This is a better use of human attention: it focuses human judgement where it adds distinctive value rather than on the mechanical assembly of risk evidence that agents can do more consistently.
IBM Concert’s change risk score [5] deserves specific attention here because it is not simply a static risk matrix. Concert evaluates the proposed change dynamically, considering the current state of the estate at the moment the change is proposed. A change to a database configuration that would be low-risk on a Sunday afternoon may be flagged as high-risk on a Tuesday at peak transaction volume, if Concert’s topology analysis shows that the affected database is currently in a degraded replication state and has two dependent services showing elevated error rates. The CAB that receives Concert’s risk score as input is receiving a live, contextualised assessment, not a template filled in at the time of change initiation without regard for current conditions.
Emergency changes represent the case where speed is paramount and normal CAB timelines are operationally unacceptable. An emergency change in response to an active P1 incident—where every minute of delay corresponds to measurable business impact—cannot wait for a scheduled CAB meeting or even an ad-hoc CAB convened in thirty minutes. The governance challenge is that emergency changes, precisely because they are applied under pressure, carry the highest risk of unintended consequences. Agents contribute most valuably at the preparation stage. An agent can assemble, in minutes, the package that an emergency CAB needs to make an informed decision: the current incident record, the Concert topology showing the blast radius of the proposed change, the change history of the affected component, any previous similar emergency changes and their outcomes, and a rollback procedure with estimated execution time. This package, which might take a human engineer an hour to assemble under pressure, arrives in the emergency CAB channel as a structured document within three minutes of the change being proposed.
The emergency CAB still approves; the agent prepares. This distinction preserves the governance intent of emergency change authorisation while removing the preparation overhead that previously made emergency governance so painful that organisations were tempted to execute first and document retrospectively—a practice that creates precisely the regulatory accountability gaps that DORA and NIS2 are designed to eliminate.
Every change record created in ServiceNow, whether for a standard, normal or emergency change, constitutes sovereign evidence in the sense described in Chapter 3: an authoritative, timestamped, attributed record of what changed, under what authority, with what justification, and with what outcome. When regulators examine the organisation’s change management practices under DORA’s ICT risk management framework or ISO 27001’s change management controls, these records are the primary evidence. Agents that create and populate these records consistently, with structured content and complete attribution, produce better evidence than human-authored records that vary in completeness and quality under operational pressure.

Problem management occupies an unusual position in the ITIL 4 practice landscape. Of the four practices discussed in this chapter, it is the most analytical, the most dependent on pattern recognition across historical data, and the most oriented toward structural improvement rather than immediate resolution. It is also the practice where AI agents add the most distinctive analytical value—and where the relationship between agent analysis and human judgement is most subtly important to design correctly.
The fundamental challenge of problem management is that root causes are often not present in any single incident record. They emerge from patterns across many incidents, patterns that span time, service boundaries, and operational domains. A database that is slowly accumulating blocking queries may be the root cause of ten separate incidents across three months, each attributed at incident closure to a proximate cause—timeouts, retries, cascading failures—without the underlying pattern being visible to any individual engineer reviewing a single incident. Problem management is, at its core, a pattern recognition problem operating across a large, noisy historical dataset.
This is precisely the analytical context in which agents operating on Concert’s knowledge graph [5] add distinctive value. Concert’s knowledge graph holds not only the current topology of the estate but the historical record of topology changes, configuration drift, signal anomalies, and incidents—all linked to the entities they concern. When an agent queries the knowledge graph with a problem investigation request, it can traverse relationships that would take a human analyst hours to assemble manually: “Show me all incidents involving this service in the past 90 days, grouped by the Concert situation classification, cross-referenced with the change records active within 48 hours of each incident, and filtered to show only those where the telemetry signature includes elevated database connection wait time.” The answer to that query, which Concert produces in seconds, may reveal that a configuration change deployed six weeks ago is the common antecedent of nine otherwise apparently unrelated incidents.
The agent’s role in reactive problem management is therefore to perform this kind of structured correlation analysis and produce a draft problem record that a human problem manager can review, validate, and act upon. The draft problem record, created in ServiceNow, contains the agent’s correlation findings, the relevant incident references, the candidate root cause with the supporting evidence from the knowledge graph, and a proposed workaround where one exists. It also contains a confidence indicator: a structured assessment of how strongly the evidence supports the proposed root cause, and what alternative hypotheses were considered and discarded. This transparency about reasoning quality is essential—problem managers who receive an agent-generated root cause without any indication of the evidence quality behind it have no basis for calibrating their confidence in the finding.
The ServiceNow Known Error Database (KEDB) is the institutional memory that problem management maintains. When an agent-identified root cause is validated by a problem manager, the corresponding known error record is created in the KEDB with the agent’s findings, the validated root cause, and the approved workaround. The KEDB record is then available to triage agents handling future incidents: if a new incident’s signature matches a known error pattern, the triage agent can immediately surface the known error and begin executing the approved workaround, reducing mean time to restore service without requiring the diagnosis to be performed again. This feedback loop—from incident investigation, through problem management, through the KEDB, back to faster incident handling—is one of the most operationally valuable patterns in the ITSM-agentic integration.
Proactive problem management leverages Concert’s continuous topology monitoring to identify conditions that have historically preceded incidents, before those conditions produce an incident. An agent running a scheduled proactive analysis might identify that a storage component’s I/O wait metric has been trending upward for three weeks, that a similar trend preceded a significant outage twelve months ago, and that the affected storage component provides critical dependencies to two tier-one regulated services. This finding warrants a proactive problem record in ServiceNow, flagging the risk for engineering attention before it becomes an incident. The agent’s analysis provides the trigger and the evidence; the human problem manager decides whether to treat it as a genuine risk and what remediation action to commission.

Service request management in a large enterprise is a high-volume, often under-resourced function. The catalogue of common requests—access provisioning, environment creation, certificate renewal, capacity adjustments, software package approvals—generates thousands of individual actions per month, most of them technically routine but requiring consistent adherence to security policies, approval workflows, and audit requirements. This is a domain where agentic automation delivers immediate, substantial operational value, and where the integration between the service catalogue, the ITSM platform, and the orchestration layer must be designed with care.
The service catalogue in an ITSM platform is not simply a menu of available services; in an agentic architecture, it is a specification of the workflows that agents execute. Each catalogue item maps to a defined watsonx Orchestrate [6] workflow: the steps, the validation checks, the approval gates, the tools that will be called, the data that will be collected, and the evidence that will be written to the service request record on completion. The catalogue item defines the human experience—what a user requests and what information they must supply—while the Orchestrate workflow defines the agent’s execution path. This separation of concerns allows the service catalogue to remain user-facing and accessible while the underlying execution becomes increasingly automated.
Consider the four most common request categories and how agents handle them.
Access provisioning is the highest-volume category in most enterprises. A user requests access to a system; the request flows into ServiceNow as a service request; an agent verifies that the requested access tier is consistent with the user’s role, that no conflicting access exists, that the appropriate approval has been obtained from the service owner, and that any required access review has been completed. If all conditions are met, the agent executes the provisioning action—calling the identity provider’s API to add the user to the appropriate group or role—and writes the evidence of the provisioning action to the service request record. The record constitutes the access review artefact that ISO 27001’s Clause A.8.2 requires: who requested what, on what basis, with what approval, and at what time. For privileged access requests—administrative access to regulated systems—the approval gate requires explicit authorisation from a named approver with appropriate authority, and the agent holds the request pending that authorisation rather than proceeding. The separation between approval and execution is enforced by the workflow, not by human discipline.
Environment creation is the second major category. A development team requests a new test environment; an agent receives the request, validates that the requested configuration is within the approved cost envelope and the permitted infrastructure footprint for the team’s cost centre, checks that the environment specification passes the IaC policy validation described in Chapter 11, and executes the Terraform or Ansible playbook that creates the environment. The service request record captures the environment identifier, the applied configuration, the IaC template version, and the policy evaluation outcome. For regulated organisations, this record also satisfies the change management obligation: environment creation is a change to the infrastructure estate, and the service request record, linked to the change record created by the agent, provides the dual-traceability that auditors expect.
Certificate renewal is operationally unglamorous but operationally critical: expired certificates are a leading cause of unnecessary service disruptions, and manual certificate management at scale is error-prone. An agent monitoring certificate expiry dates—reading from the Concert CMDB or from a dedicated certificate management system—identifies certificates approaching expiry and initiates the renewal workflow automatically. The service request record provides the audit trail: the certificate identifier, the renewal date, the new expiry, the validation that the renewed certificate was correctly installed and the service verified as operational. For regulated entities subject to NIS2’s cryptography requirements under Article 21(2)(viii), this audit trail is the evidence that cryptographic asset management is being performed in a controlled and documented way.
Capacity adjustments represent a class of request that spans the boundary between service requests and change management: a team requests additional compute capacity for an anticipated load event, or a monitoring alert indicates that current capacity will be insufficient within a defined horizon. An agent evaluates the request against the current resource allocation, the available headroom in the sovereign zone where the service runs, the cost implications, and the risk profile of the proposed scaling action. For requests within pre-approved bounds—scaling within a range that the service’s capacity model has validated—the agent executes without additional human approval and records the action in the service request record. For requests outside those bounds—scaling that would approach zone capacity limits, or scaling that would require infrastructure changes in another cloud provider’s region—the agent creates a normal change record and routes it through the CAB process described in Section 19.4.
SLA tracking for agent-executed service requests requires particular attention to the velocity difference between automated and manual execution. When a human service desk agent handles requests, SLA compliance is measured in hours and is an operational challenge. When an AI agent handles requests, execution typically completes in minutes, but the SLA clock starts from the moment the request is submitted, not from the moment the agent begins work. SLA breaches in an automated environment are therefore most commonly caused not by slow execution but by approval wait times, system availability windows, or policy evaluation delays. The SLA monitoring component of the Orchestrate workflow must track all of these contributions to the elapsed time, surface pending approvals that are at risk of breaching the SLA, and escalate to the service desk manager before the breach occurs rather than after.

The ITSM integration described in the preceding sections addresses the operational lifecycle: incidents, changes, problems, and service requests managed through ServiceNow. But many operational actions have a consequence that extends into the development domain. Infrastructure problems require code changes. Security vulnerabilities require dependency updates. Capacity constraints require architectural modifications. In a DevOps organisation, these consequences must cross the boundary between operations and development, and they must do so in a way that preserves the traceability that both domains require. Jira and GitHub are the canonical platforms on which development work is managed and version-controlled, and their integration into the agentic operations workflow is an architectural necessity rather than an optional enhancement.
The primary integration pattern is agent-generated work items. When an operational agent determines that a resolution to an incident or problem requires a code or configuration change that falls within the development domain—a dependency upgrade, a configuration parameter that must be managed through version-controlled IaC rather than runtime modification, a performance optimisation that requires application code changes—the agent creates a Jira ticket representing the required work. The ticket is not a bare request; it is enriched with the operational context that makes it immediately actionable by the development team. The Concert situation identifier, the ServiceNow incident or problem record reference, the relevant telemetry evidence, and the agent’s analysis of the technical change required are all included as structured fields or as linked artefacts. A developer picking up the Jira ticket does not need to investigate the operational context from scratch; the agent has assembled it.
The Jira ticket is linked to the originating ServiceNow record through a cross-system identifier, typically the Concert situation identifier, that serves as the common key across all records in the operational event. This cross-referencing is not merely administrative tidiness; it is the mechanism by which the unified operational record is constructed. When an auditor or a post-incident review team asks “what happened to this incident from end to end?”, the answer can be traced from the Concert situation, through the ServiceNow incident record, to the Jira ticket that represents the longer-term resolution work, and ultimately to the GitHub pull request and commit that implemented the fix.
The second integration pattern is agent-generated pull requests. For infrastructure changes that are fully within the IaC model—configuration adjustments, security group rule modifications, resource tag updates—an agent can generate the corresponding code change directly in the Git repository and open a pull request. The pull request description is populated with the operational justification: the Concert situation, the ServiceNow change record number, the specific policy or operational condition that the change addresses, and the validation evidence that demonstrates the change is correct. The pull request enters the standard code review and CI/CD pipeline, providing the same engineering rigour that human-authored changes receive, but without requiring a human engineer to translate the operational requirement into code.
This pattern is particularly valuable for security-related changes. When Concert identifies a configuration drift—a security group rule that has been modified away from the IaC-defined baseline, a storage bucket that has acquired a public access permission that the policy model does not allow—an agent can generate the corrective IaC change and open a pull request to reinstate the correct configuration. The change is then reviewed by a security engineer or an automated policy check before being merged, ensuring that the correction does not introduce new issues. The pull request provides the evidence that the drift was identified, that a correction was proposed and reviewed, and that it was merged and deployed. This evidence chain satisfies ISO 27001’s configuration management requirements and, in environments subject to DORA’s ICT risk management obligations, contributes to the evidence of continuous control effectiveness.
The unified operational record—the collection of linked artefacts across Concert, ServiceNow, Jira, and GitHub that together tell the story of an operational event from detection to full resolution—is the most complete expression of operational traceability in a modern DevOps enterprise. It is not assembled manually; it is constructed incrementally by agents that write cross-references as they create each artefact. By the time an operational event is fully resolved, the record exists whether or not anyone has thought to build it, because the integration patterns ensure that it is built automatically as a side effect of the agents doing their primary work.

The integration architecture for Jira and GitHub follows the same agency-through-API pattern used for ServiceNow, but with two additional considerations that reflect the development domain’s different governance culture. First, Jira tickets created by agents are flagged with an “agent-created” label and include a structured provenance section that identifies the creating agent, the Concert situation that triggered creation, and the operational record that contains the supporting evidence. This transparency is important for development teams who need to distinguish between user-reported feature requests, manually created engineering tasks, and operationally-motivated work items generated by agents. Second, pull requests created by agents are submitted from a designated service account rather than from a human developer’s account, and they are tagged with a label that triggers an additional review step in the CI/CD pipeline—a step that specifically checks that the change is consistent with the operational justification provided in the pull request description. This additional gate ensures that agent-generated code changes receive appropriate scrutiny even when they are routine from the agent’s perspective.
The preceding sections have described how agent-assisted ITSM integration creates a rich, interconnected body of operational records. This section addresses the explicit regulatory and compliance value of those records, and the mechanisms by which evidence is extracted from them for audit purposes.
DORA’s Article 17 requires financial entities to establish and implement incident management processes that include the capability to classify incidents by type and severity, to record all relevant information and developments, and to establish formal closure requirements including post-incident analysis for major incidents [1]. The agent-assisted incident lifecycle described in Section 19.3 satisfies each of these requirements structurally: classification is performed by Concert’s impact scoring and reflected in the ServiceNow incident priority; all agent actions are recorded as attributed activity entries on the incident record; and closure is gated on the agent-generated post-incident summary. There is no additional administrative effort required to produce the DORA incident record; the record is the incident record, and it is created and maintained throughout the incident by the agents and the ITSM integration layer working together.
NIS2’s notification obligations under Article 23 require early warnings within 24 hours of awareness, full incident notifications within 72 hours, and final reports within one month [2]. These obligations are tracked as compliance timers on the ServiceNow incident record, managed by the communication agent described in Section 19.3. The 24-hour early warning, when due, is drafted by the communication agent using the structured content available in the incident record at that point—severity, initial impact assessment, indicators of compromise where available—and routed for human review before submission. The 72-hour notification is similarly pre-drafted and enriched by whatever additional agent-generated findings are available by that point in the investigation. The one-month final report draws on the complete incident record, the linked problem record, and any associated change records, compiled by a closure agent into the structured format required by the relevant national CSIRT. In each case, the agent’s role is to reduce the manual effort of compliance reporting to a review and approval action rather than a drafting exercise.
ISO 27001:2022’s change management requirements, expressed in Clause A.8.32, specify that changes to information processing facilities and systems shall be controlled, with documented procedures covering planning, testing, authorisation, implementation, and review [3]. The change records described in Section 19.4, created and maintained by agents for all three categories of change, satisfy this requirement directly. The structured change record includes the planning artefact (the risk assessment and change model reference), the testing evidence (pre-change state validation), the authorisation record (human approval or standard change model pre-approval reference), the implementation evidence (agent execution log with before-and-after state), and the post-implementation review (Concert’s post-change health monitoring result). An ISO 27001 auditor examining the organisation’s change management evidence finds, in a well-configured agentic environment, change records that are more complete and more consistently structured than those produced by manual processes, because agents do not omit fields under time pressure or use free-text narratives where structured data is required.
The extraction of compliance evidence from the interconnected ITSM and Concert records is a function that can itself be automated. A compliance evidence extraction agent, running on a scheduled basis or triggered by an audit request, can query the ServiceNow API for all incident records in a specified time window meeting defined classification criteria, retrieve the associated Concert situation identifiers and knowledge graph context, compile the set of records into a structured evidence package, and deliver it to the compliance team in a format ready for audit presentation. This does not replace the judgement required to assess whether the evidence is adequate; it eliminates the manual effort of assembling it, which in traditional operations environments can occupy compliance teams for weeks before each regulatory submission.
The governance architecture must also address the meta-level compliance question: how does the organisation demonstrate that the agents themselves are operating within their defined bounds? Watsonx.governance [7] provides the audit capability here. Every agent invocation that results in a ITSM action is logged in the watsonx.governance audit trail, including the agent identity, the policy evaluation outcome from OPA, the human approval reference where one exists, and the artefact identifier of the ITSM record that was created or updated. This audit trail is separate from and complementary to the ITSM records themselves: the ITSM records demonstrate what was done to the infrastructure; the governance audit trail demonstrates that the agent’s authority to do it was properly constituted and that its actions were within the bounds of that authority.
The combination of these two audit trails—the ITSM operational record and the governance agent behaviour record—provides the dual-layer evidence that the most demanding regulatory frameworks require. DORA’s requirements for ICT risk management framework adequacy, NIS2’s requirements for documented cybersecurity measures, and ISO 27001’s management system evidence requirements are all satisfied through records that the agentic ITSM integration produces as a natural consequence of operating correctly, without separate evidence production effort.

ITSM processes are not in tension with agentic operations; they are the institutional accountability layer that makes automated operations governable, traceable, and acceptable to regulators. Agents must operate through ITSM processes, not around them, if automated operations are to satisfy DORA, NIS2, and ISO 27001 obligations.
The four ITIL 4 practices most fundamentally affected by agent participation—incident management, change enablement, problem management, and service request management—all require ITSM tooling capable of consuming machine-generated, structured content at scale and supporting mixed human-agent workflows within a single operational record.
Agent-assisted incident management produces consistently richer incident records than manual processes, because agents log every investigative action as an attributed entry on the ITSM record in real time, without the selective documentation that occurs when human engineers are under pressure. Concert’s situation identifier, written into the ServiceNow incident record at creation, anchors the unified operational record that spans both platforms.
Change management in an agentic environment is differentiated by change tier: standard changes are executed autonomously against pre-approved models with OPA policy validation; normal changes have their risk assessment generated by Concert’s dynamic change risk scoring for CAB review; emergency changes are prepared by agents for fast-track human authorisation. All three tiers produce ServiceNow change records that constitute sovereign evidence under applicable regulatory frameworks.
Problem management is the practice where agents add the most distinctive analytical value, because Concert’s knowledge graph enables structured correlation analysis across incident histories, topology changes, and telemetry patterns at a scale and speed that is practically unreachable by human analysts working against the same datasets.
The unified operational record—Concert situation linked to ServiceNow incident and change records, linked to Jira development tickets, linked to GitHub pull requests through a common cross-reference key—is constructed automatically by agents as a side effect of their primary work, providing complete operational traceability from detection to resolution without requiring any separate record assembly.
ITSM-integrated agentic operations satisfies DORA Article 17 incident records, NIS2 Article 23 notification obligations, and ISO 27001:2022 Clause A.8.32 change management evidence requirements through records that agents create and maintain as a natural consequence of operating correctly, eliminating the gap between operational practice and compliance evidence production that is endemic to manual operations.
This chapter has examined how agentic operations integrates with ITSM disciplines to produce governable, traceable, and regulatory-compliant operational outcomes. The integration patterns described here—Concert situations linked to ServiceNow incidents, change requests, and problem records; agents creating Jira tickets and GitHub pull requests as operational events cross the development boundary; compliance evidence automatically extracted from the interconnected record set—depend on agents that have access to the operational context required to perform these tasks correctly. Agents must know what a given service does, what its topology looks like, what its operational history has been, and what patterns of behaviour have preceded incidents in the past.
This contextual knowledge has, throughout this part of the book, been attributed primarily to Concert’s topology model and its knowledge graph. But there is a category of operational knowledge that Concert’s continuously discovered topology cannot capture on its own: the accumulated human understanding of systems, the documented architectural decisions, the runbook annotations, the post-incident analysis findings, and the domain expertise that exists in technical documentation, architecture records, and the institutional memory of operations teams. Chapter 20 examines how this knowledge is captured, indexed, and made available to agents through knowledge-augmented generation—the retrieval of relevant documented knowledge to supplement Concert’s structural topology with the semantic and contextual understanding that agents need to reason correctly about the most complex operational situations.
[1] European Parliament and Council, “Regulation (EU) 2022/2554 of the European Parliament and of the Council on digital operational resilience for the financial sector (DORA),” Official Journal of the European Union, L 333, pp. 1–79, 27 Dec. 2022. [Online]. Available: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32022R2554
[2] European Parliament and Council, “Directive (EU) 2022/2555 of the European Parliament and of the Council on measures for a high common level of cybersecurity across the Union (NIS2 Directive),” Official Journal of the European Union, L 333, pp. 80–152, 27 Dec. 2022. [Online]. Available: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32022L2555
[3] International Organization for Standardization, ISO/IEC 27001:2022 — Information Security, Cybersecurity and Privacy Protection — Information Security Management Systems — Requirements, Geneva: ISO, 2022. [Online]. Available: https://www.iso.org/standard/27001
[4] Axelos, ITIL 4 Foundation: ITIL 4 Edition, London: TSO, 2019. [Online]. Available: https://www.axelos.com/certifications/itil-service-management
[5] IBM, “IBM Concert: Introduction and overview,” IBM Documentation, 2024. [Online]. Available: https://www.ibm.com/docs/en/concert
[6] IBM, “IBM watsonx Orchestrate: Agent and tool configuration,” IBM Documentation, 2024. [Online]. Available: https://www.ibm.com/docs/en/watsonx/watson-orchestrate
[7] IBM, “IBM watsonx.governance: Model risk governance and AI lifecycle management,” IBM Documentation, 2024. [Online]. Available: https://www.ibm.com/docs/en/watsonx/saas?topic=overview-watsonxgovernance
[8] Atlassian, Atlassian Incident Management Handbook, Atlassian Corporation, 2023. [Online]. Available: https://www.atlassian.com/incident-management
[9] ServiceNow, “ServiceNow IT Service Management documentation,” ServiceNow Documentation, 2024. [Online]. Available: https://docs.servicenow.com/bundle/washingtondc-it-service-management
Sovereign Cloud Operations: AI-Driven Management of Sovereign Estates
© 2026 by Alan Hamilton
is licensed under CC BY-SA 4.0