Guide
AI for IT operations: what agents can safely take on
AI for IT operations means using AI models and agents on the work of running IT: sorting and routing tickets, answering staff questions from your knowledge base, grouping alerts, summarising incidents and running known fixes from runbooks. What an agent can safely take on depends on what it is allowed to change. Work that only reads and suggests, such as triage suggestions and incident summaries, can start early at little risk. Work that changes systems, such as granting access, restarting services or changing configuration, needs controls first: the agent's own identity with the least access that does the job, deterministic checks on every action it proposes, your existing change control applied to it as to a person, human approval for high-impact actions, and a log of everything it ran. Hand work over in that order, and move each stage forward only on evidence from real tickets.
AI for IT operations. The use of AI models and agents to triage, answer, summarise and resolve IT service and operations work, within controls on what they are allowed to change.
Checked . Vendor features were read in Atlassian's own documentation and in articles by ServiceNow employees on ServiceNow's community site, and the NIST, OWASP and Google guidance on the publishers' own pages, on this date. Vendors change their AI products and documentation often, so confirm the details on the sources below before you rely on them. This guide is not legal or compliance advice.
Which IT operations work AI can take on, and what each job lets it change
I sort IT operations work by one question: what is the AI allowed to change? A model that suggests a category for a ticket can be wrong at almost no cost, because an analyst sees the suggestion before it matters. A model that restarts a service or grants access can be wrong at three in the morning with nobody watching. The table ranks the common jobs by that exposure, lowest first.
| Dimension | What the AI does | What it can change | Control before it starts |
|---|---|---|---|
| Ticket triage and routing | Suggests a category, priority and resolver group for each new ticket | Nothing while an analyst accepts each suggestion; the routing itself once it assigns tickets on its own | A test set of past tickets with the group that finally resolved each one, and a way for analysts to correct it |
| Knowledge answers for staff | Answers staff questions from your own help articles, in the portal or in chat | What staff believe and do next, so a wrong answer spreads | Answers drawn only from current articles, each citing its article, with a hand-off to a person |
| Incident summaries and reviews | Drafts incident summaries, timelines and the first version of a post-incident review | Nothing directly; the record, once someone accepts the draft | A named person who edits and signs off every review |
| Alert grouping | Clusters related alerts so on-call staff see one problem rather than fifty | What people look at, so a wrong group can hide a real incident | A count of the groups people later split by hand, and a way to switch grouping off |
| Access requests | Checks a request against policy and prepares the grant for approval | Who can reach which system | The agent's own identity, an approver's sign-off, and an audit record of every grant |
| Runbook fixes for known faults | Runs a pre-approved fix, such as restarting a service or clearing a queue | Production systems | A dry run, a limit on how often it may act, a way back, and the fix handled as a standard change |
| Configuration and infrastructure changes | Proposes a change to configuration, firewall rules or infrastructure code | Production systems, often more than one at once | Your full change process with a person approving, and an agent that never applies its own change |
The first three rows are where I would start almost any IT team. They save analyst time from the first week, and when they go wrong, they go wrong where people can see it. The last three are where the savings get larger, and where every control in the sections below has to exist first.
Vendors sell part of this as AIOps: machine learning over monitoring and event data, to group alerts, spot anomalies and suggest likely causes. The label covers the alert-grouping row and some of the incident work. It says little about the rows that change systems, and those are the ones a CIO ends up answering for.
Start with toil, and check whether it needs a model at all
Google's site reliability engineering book defines toil as work tied to running a production service that tends to be manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows, and Google's SRE organisation aims to keep it below half of each SRE's time [1]. That definition is a good filter for where AI goes first, because toil is the work people are glad to give away.
A good share of IT toil needs no model. When a client's IT team was over capacity, we built an identity provisioning system for more than 1,000 users on explicit rules: access follows the HR record, each run applies only the difference between intended and actual state, changes can be inspected before they land, and a person is matched across systems by rule rather than by fuzzy matching, because a wrong match grants the wrong person access. If a job can be written down as rules, write the rules. Save the model for the part that needs judgement across messy text, such as working out what a badly written ticket is actually asking for.
The same test applies to automations that already run on a no-code platform. If one is straining, our guide to when Zapier is not enough sets out the signs and what to try before rebuilding.
The controls an agent needs before it touches a system
An agent that can act on your estate is a new operator, and it should meet the bar you set for a new engineer with production access, plus a few checks that exist because it is a model. OWASP calls the failure these controls prevent excessive agency: damaging actions taken in response to unexpected, ambiguous or manipulated model output, caused by too much functionality, too many permissions or too much autonomy [2].
Where NIST's security and privacy control catalogue, SP 800-53, already words a control well, I cite it [3]. For a private company it is guidance rather than law. I use it because its wording is precise.
- Its own identity, with least access
- The agent acts as a named service identity, never a shared admin account and never a person's borrowed credentials, with only the permissions its job needs. NIST's least privilege control allows only the access needed to accomplish assigned tasks, for users and for processes acting on their behalf [3]. OWASP recommends limiting both the tools an agent may call and the permissions those tools hold [2].
- Read first, write later
- Give it read access for the first stage and write access job by job, as the evidence supports it. A triage agent does not need to close tickets to be useful.
- Change control applies to the agent
- A restart, a rule change or an access grant made by an agent is a change like any other. SP 800-53 asks you to decide which changes are controlled, to review and approve them with their security and privacy impact considered, to record the decisions and keep the records, and to analyse a change's security and privacy impact before it is made [3]. Route the agent's changes through the process you already run. Pre-approved standard changes are the natural first place for an agent to act, because a person approved that defined, low-risk fix in advance.
- Checks the agent cannot talk past
- Every action it proposes passes deterministic checks before it runs: the host is in scope, the action is on the allowed list, the time is inside the change window, the rate limit has not been reached. That is the pattern where the agent proposes and a verifier gates, applied to operations. The check is ordinary code, so no clever wording in a ticket can argue with it.
- A person for high-impact actions
- OWASP recommends human approval for high-impact actions, and authorisation enforced in the downstream system rather than left to the model [2]. Show the approver the specific action and the evidence behind it, not the whole conversation.
- Every privileged action logged
- Log what the agent read, what it proposed, which checks ran, who approved and what changed. NIST's catalogue asks for the execution of privileged functions to be logged [3]. You will want that record the first time someone asks why a server restarted overnight.
- Ticket text is untrusted input
- Anyone who can raise a ticket or email the service desk can write text the agent will read. Treat all of it as untrusted: the agent's permissions decide what it can do, and nothing written in a ticket can widen them.
- A limit and an off switch
- Cap how many actions it takes per hour and per system, and give the on-call engineer one switch that stops it. An agent looping on a flapping alert should hit its limit long before it breaks what it was meant to fix.
What your service management platform already ships
If you run ServiceNow or Jira Service Management, some of this arrives with the licence, and it is worth knowing what before anyone proposes a build. I describe each from the vendor's own documentation or, for ServiceNow, articles its employees publish on its community site, read on the date at the top of this page. This is not a ranking.
ServiceNow lets each tool an AI agent uses run in a supervised mode, where the agent asks the user to confirm its reasoning, or an autonomous mode, where the user sees only the output [4]. An article on access controls by a ServiceNow employee recommends supervised mode for agents that can take sensitive or critical actions. It describes two identities an agent can run as: a "dynamic user", the default, which inherits the permissions of the person who invoked it, or a dedicated AI user with its own preconfigured roles. It adds that role masking, introduced in the Zurich Patch 4 release, limits an agent to a subset of the invoking user's roles [5].
Atlassian lists AI features in Jira Service Management that include a virtual agent that answers questions from your linked knowledge base, grouping of related alerts, drafted post-incident reviews, incidents created from alerts, and suggested actions on service requests and incidents [6].
My view is to use the platform's own AI for work that lives entirely inside the platform, such as summaries, suggested routing and knowledge answers, because it already holds your tickets and your permission model. Build a layer of your own when the work crosses systems the platform does not own, such as your cloud accounts, your identity provider and your deployment pipeline, or when you need checks the platform cannot express. Either way, the controls above still have to exist. A vendor's supervised mode is one of them, not all of them.
How to test an IT operations agent before it acts
Your ticket history is a test set you already own. Take a few months of closed tickets and replay them: what would the agent have suggested, and does it match what actually happened?
Score triage against the group that finally resolved each ticket, not the one it was first sent to, because a ticket that bounced between teams was misrouted at the start. Score knowledge answers on whether each one cites a current article and whether that article supports the answer. Score alert grouping on the groups your on-call staff later split by hand, because a wrong merge can bury a real incident inside a noisy group.
Then run in shadow. The agent makes its suggestion or prepares its action alongside the analyst, who does the work as normal, and you compare the two. Agree the pass mark before anyone sees a score, move a job to the next stage only when it clears that mark on real tickets, and rerun the whole set whenever the prompt, the model, the tools or the knowledge base change.
Measure the service as well as the model. Reassignment rate, reopened tickets, time to resolve and the number of escalations tell you whether staff and users are better off, and that is the number a budget holder will ask for. If your risk team works to NIST's AI Risk Management Framework, which NIST describes as intended for voluntary use and is currently revising [7], these results are the evidence to take to it.
A rollout order that keeps change control intact
Each stage earns the next. The table sets out the order I would use and what should be true before a team moves on.
| Dimension | What the agent does | Move on when |
|---|---|---|
| Stage 1: suggest | Suggests routing and drafts replies and summaries; people act on them | Suggestions clear the pass mark you agreed on your replayed tickets, and analysts are using them |
| Stage 2: answer staff | Answers staff questions from your knowledge base, with a hand-off to a person | Answers cite current articles, and reopened or escalated requests are no higher than your baseline |
| Stage 3: route on its own | Assigns tickets itself where its confidence is high and leaves the rest to people | Its tickets are reassigned no more often than tickets people route |
| Stage 4: run standard changes | Runs pre-approved runbook fixes, with a dry run, limits and a log | Every run in the period has a clean record, and the off switch has been tested |
| Stage 5: propose other changes | Prepares other changes for your normal approval process, and never approves its own | It does not move on: a person keeps approving |
Security incidents sit outside this ladder. An agent can gather evidence and draft the timeline, but the decisions in a security incident belong to the people your incident response plan names. NIST's incident response guidance, revised in April 2025 and now written as a profile of its Cybersecurity Framework 2.0, is a sound reference for that plan [8].
The order matters more than the speed. I would rather a team spent an extra month at stage three than skipped to stage four and spent the next quarter rebuilding trust after one bad night.
Where 1AYM fits
This is the work our AI integration and automation services do: AI built into the tools your IT team already runs, with the dry runs, rollback, audit logs and identity handling that let it write to live systems, and with agentic workflow design supplying the checks on each action. We would start with one queue or one runbook, test it against your ticket history, and take it through the stages above. The provisioning system described earlier is the standard we build to: deterministic where it can be, reversible and logged.
If you are not yet sure which jobs to start with, a fixed-scope AI Opportunity & Feasibility Sprint settles that and sets out what the build involves. A fixed-scope build can start within a day of the scope being signed, and if you already have a scoped job and need people to deliver it, we can resource it on contract from the collective of associates who work with 1AYM, held to the same standard. The first step is a 30-minute call about the queue or runbook you have in mind.
For engineers: identity, tools, gates, logging and tests for an IT operations agent
The controls above, as they land in configuration. Each one can be checked in the estate rather than taken on trust.
- Service identity
- One identity per agent and per environment, issued by your identity provider, with scoped roles and short-lived credentials. No shared admin accounts, no personal tokens, and no standing access to production from a development environment.
- Narrow tools
- Expose specific operations (restart this service, clear this queue, read this log) rather than a shell or a general API client. OWASP's mitigations for excessive agency include avoiding open-ended extensions in favour of granular ones [2].
- Authorisation downstream
- Enforce what the agent may do in the target system's own permissions, so a manipulated prompt cannot widen them [2].
- Change records
- Every write opens or references a change record linked to the affected configuration item, so the agent's changes appear in the same reports and reviews as everyone else's, and the records are kept [3].
- Dry run and diff
- Each action computes the intended change and shows it before applying it. Applied changes are idempotent, so a retry cannot compound a mistake.
- Limits and a kill switch
- Per-system and per-hour action caps, a circuit breaker on repeated failures, and one feature flag the on-call engineer can flip without a deploy.
- Untrusted text
- Ticket bodies, email and log lines reach the model as data, never as instructions to the tool layer. Test with tickets written to steer the agent.
- Tracing
- Record the input, the retrieved articles, the model output, the checks run, the approver and the result for every action, and send it to the log store your security team already monitors. Log privileged functions explicitly [3].
- Replay suite in CI
- A labelled set of historical tickets and alerts, awkward ones included, that runs on every change to the prompt, model, tools or knowledge base, against thresholds agreed before the first run.
Sources
- [1]Google, Site Reliability Engineering, chapter 5: Eliminating Toil (Vivek Rau, edited by Betsy Beyer), read 29 September 2026
- [2]OWASP Gen AI Security Project, LLM06:2025 Excessive Agency, read 29 September 2026
- [3]NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations: controls AC-6, AC-6(9), CM-3 and CM-4 as published in Release 5.2.0 (27 August 2025), read 29 September 2026
- [4]ServiceNow Community, Create your own AI Agent! A walkthrough on creating an AI Agent using AI Agent Studio (ServiceNow employee, 17 March 2025), read 29 September 2026
- [5]ServiceNow Community, Latest access control security enhancements for AI Agents and Skill Kit [Updated Zurich Patch 4] (ServiceNow employee, first posted 9 September 2025), read 29 September 2026
- [6]Atlassian Support, AI features in Jira Service Management, read 29 September 2026
- [7]NIST, AI Risk Management Framework (AI RMF 1.0 released 26 January 2023), read 29 September 2026
- [8]NIST SP 800-61 Rev. 3, Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile (April 2025), read 29 September 2026
Frequently asked questions
What is AIOps, and is it the same as AI for IT operations?
AIOps is the name vendors use for applying machine learning to monitoring and event data: grouping alerts, spotting anomalies and suggesting likely causes. AI for IT operations is wider. It also covers the service desk, knowledge answers and agents that carry out fixes, which is where most of the control questions sit.
Can an AI agent approve its own changes?
No. Treat the agent as the requester, never the approver. A pre-approved standard change is not an exception: a person approved that defined fix in advance, and the agent only runs it.
Should we use ServiceNow's or Atlassian's AI, or build our own?
Use the platform's own AI for work that stays inside it, such as summaries, suggested routing and knowledge answers. Build your own layer when the work crosses systems the platform does not own, or needs checks it cannot express. The two can sit side by side, and the controls on this page apply to both.
How long before an agent can resolve tickets without a person?
As long as the evidence takes. Move a job forward when it clears the pass mark you agreed on replayed and shadowed tickets, and keep a person approving anything beyond pre-approved standard changes. A date set in advance tends to become the date the controls get skipped.
Further
- AI integration and automation services · The service this guide leads to: AI inside the systems your IT team already runs, with audit logs and rollback.
- Agentic workflow design · The deterministic checks that decide which of an agent's actions proceed.
- AI agents for business · Where agents fit across a business, beyond IT, and where they do not.
- RPA vs AI automation · When to keep, extend or replace the bots already running.
We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.
Last reviewed · 1AYM