Guide

AI workflow automation: which processes to automate first

Automate first the process that scores well on value, feasibility and risk at once. Value means it runs often, takes real time per case and its errors cost money today. Feasibility means its inputs sit in systems you can connect to, its rules can mostly be written down, and past cases with known outcomes exist to test against. Low risk means a wrong output is caught or can be undone before it harms anyone. The most valuable process is often not the right first one, because the first build also puts the connections, logging and tests in place for everything after it. Use fixed steps wherever the path can be drawn in advance, with an AI step only where the input needs reading or judging, and keep agents for work whose steps change from case to case.

AI workflow automation. Running a business process across systems with software that follows fixed steps, uses an AI model for the parts that need reading or judgement, and sends exceptions to a person.

1AYM is an OpenAI Select Partner. Its founder holds personal Claude certifications. 1AYM is not an Anthropic partner. This guide cites both vendors' published guidance on agents and ranks neither. 1AYM also builds AI integration and automation, one of the routes this guide describes, so weigh our view with that in mind.

Checked . The vendor guidance and the UK government, ICO and OWASP sources this page cites were read on the publishers' own sites on 29 September 2026. The worked examples are fictional: their processes and scores show the method and are not client results. Where the page touches data protection law it summarises published guidance and is not legal advice.

Three shapes of AI workflow automation

Ask three suppliers to quote for AI workflow automation and you can get three different builds under the same name. The difference decides the cost, the risk and how the thing gets tested, so settle which shape a process would take before you score it.

Fixed steps
Software follows a path drawn in advance: an invoice arrives, it is matched to the order, posted and sent for approval. No model is involved. Where the inputs are already structured this is often the whole answer, and it is the cheapest shape to run and to test.
Fixed steps with an AI step
The same drawn path, with a model inside the one or two boxes that need reading or judgement, such as pulling fields from a PDF, classifying an email or drafting a reply. The path, the permissions and the checks around the model stay fixed.
An agent
A model decides which steps to take, in what order and with which tools, until it judges the task done. It earns its place when the path differs from case to case, and it is harder to test and costs more per case.

This guide is about choosing and sequencing the processes. The build itself, with audit logs, dry runs and rollback on every write to your systems, is what our AI integration and automation services cover.

Who builds it is a separate choice. An AI automation agency usually builds on a no-code platform and often runs the result for a monthly fee, while a consultancy starts from the business problem. We set out the difference between an AI agency and an AI consultancy, including what you own when the contract ends, in a separate guide.

Score each candidate on value, feasibility and risk

Start the long list with the people who do the work, not with a vendor's list of use cases. The UK government's AI Playbook says the choice of use case must be led by business and user needs, pain points and inefficiencies, not by what the technology can do. That holds well outside government.

Then score each candidate from 1 to 3 on the nine criteria below, using measured figures wherever you have them. Risk is scored the other way round from the other two, so that 3 is always the better score: a 3 on risk means a wrong output is cheap, caught early and easy to undo.

Nine criteria for scoring a process for AI workflow automation, 1 to 3 each, with risk scored so that 3 is safest
DimensionWhat to look atScores 3 whenScores 1 when
Value: volumeCases a month, counted from the system that logs them rather than estimatedHundreds or thousands a month, every monthA handful a month, or one burst a year
Value: time per caseHandling time, from system timestamps where they existTens of minutes of skilled time per caseA minute or two
Value: cost of errors todayRework, write-offs, penalties and complaints the current process causesErrors today cost real money or customersErrors are rare and cheap to put right
Feasibility: system accessWhere the inputs arrive and where the outputs must be writtenEvery system involved has a usable API or a supported exportThe work lives in email threads, personal spreadsheets or a screen with no API
Feasibility: rules and past casesHow much of the decision can be written down, and whether past cases with known outcomes exist to test againstThe rules are written or can be, and hundreds of past cases have recorded outcomesNobody can say how the decision is made, and outcomes are not recorded
Feasibility: stabilityHow often the process, its forms or its systems changeStable for at least the next yearBeing redesigned, or a system it depends on is about to be replaced
Risk: cost of a wrong outputWhat a wrong answer does if it gets throughAn internal record someone correctsMoney leaves, a customer is told something false, or a person is refused something
Risk: caught and reversibleWhether a check catches the error before it causes harm, and whether the action can be undoneA rule or a reconciliation catches it, and the action can be reversedNothing checks it, and the action cannot be undone, such as a payment or a sent letter
Risk: decisions about peopleWhether the output decides something significant about an individualNo individual is affectedIt decides someone's access to money, work, services or care

Keep the three groups as three totals rather than adding them into one. A process scoring 9 on value and 3 on risk needs a different conversation from one scoring 6 on everything, and a single number hides the difference.

The scores rank candidates against each other. They are not a business case, and before a build is approved the chosen process still needs a measured baseline and a proper estimate of the return.

Four findings that stop a candidate before scoring

Some findings end the conversation for now, whatever the scores say. I check for these first, because any one of them can sink a project that looks excellent on paper.

No named owner
Nobody who runs the process today has the time to explain it, supply past cases and check the results. Without that person, the build guesses at the rules.
A process nobody agrees on
If the process is being redesigned, or the people in it disagree about how it should run, automating it makes a bad process faster. Settle the process first.
No way to check an output
No past cases with known right answers, and no rule that can confirm an output is correct. Build that test set before anything else; it is cheaper than finding the errors in production.
Software alone deciding about a person
Under UK GDPR as amended by the Data (Use and Access) Act 2025, a significant decision based solely on automated processing needs safeguards: telling people about the decision, letting them make representations and challenge it, and letting them obtain human intervention. The ICO notes that the wider choice of lawful bases for these decisions does not extend to special category data. Keep a person making the decision, or take advice before designing it any other way. This is not legal advice.

What goes first, what goes next

The process with the highest value score is often the wrong first choice. A first build has two jobs: to return something, and to lay the groundwork every later automation reuses, meaning the connections to your systems, the logging, the habit of testing against past cases and the routes to a reviewer. A process that is easy to connect and safe to get wrong does the second job well, even when its value is only middling.

So I would sequence the list in rounds, re-scoring the rest after each one, because the first build changes what is feasible for everything after it.

First
Feasibility and risk totals of 7 or more, and value of at least 6: something that runs every day, where a wrong output is caught before it matters. The aim is a working system in production with its logs and tests, not a showcase.
Next
Candidates that reuse the systems the first build connected, including ones with messier inputs or more risk. Take on more risk only once the first automation has run for a full business cycle and its error rate has been measured.
Later, with checks in front
High value and high risk, such as anything that pays out money or decides something for a customer. These go last, with a deterministic check between the AI step and the write, and a person on every consequential case.
Not yet
Anything with a feasibility total of 5 or less. Improve the data, the system access or the process itself, and score it again in a quarter.

When a workflow needs an agent, and when fixed steps are enough

Anthropic's engineering guidance draws the line clearly. In its terms, workflows are systems where models and tools are orchestrated through predefined code paths, and agents are systems where the model directs its own process and tool use. It recommends finding the simplest solution possible and adding complexity only when needed, and it notes that agents trade latency and cost for better task performance, with the potential for compounding errors.

OpenAI's practical guide to building agents reaches the same place from the other side. It suggests prioritising workflows that have resisted automation: decisions full of judgement and exceptions, rule sets that have become costly to maintain, and heavy reliance on unstructured data such as documents and conversation. If a use case does not clearly meet those criteria, it says, a deterministic solution may suffice.

My own test is simpler. If the people doing the work can draw the flowchart, build the flowchart and put a model in the boxes that need reading or judgement. Reach for an agent when the number and order of steps change from case to case, and a drawn path would need dozens of branches to cover what a person does. In the first year of an automation programme I would expect many more of the first kind than the second.

Fixed steps with an AI step inside, compared with an agent, on how each is built, tested and run
DimensionFixed steps with an AI stepAn agent
The path through the workDrawn in advance; the model fills in one stepChosen by the model while it runs
TestingEach AI step is tested on its own against past cases, and the path is tested like any other softwareWhole runs have to be tested, because the path itself can go wrong
Cost per casePredictable: a known number of model callsVaries with how many steps the model takes
How it failsOne step gives a wrong output, which a check after it can catchA wrong early choice can carry through every later step
PermissionsEach step gets only the access that step needsThe agent's tools set the limit, so each tool needs the narrowest access possible
Typical fitsInvoice coding, email triage, document extraction, drafting replies for a person to sendRequests and investigations whose steps differ each time, within a small set of tools

Three candidates, scored (fictional)

These three processes, from three different organisations, are invented to show the method. The organisations, volumes and scores are illustrative, and none of them is a client result or a benchmark. Each is scored from 1 to 3 on the nine criteria above, with risk scored so that 3 is safest.

Fictional worked examples: three processes scored for AI workflow automation
DimensionSupplier invoice codingDelivery-change requestsHardship payment plans
Volume332
Time per case213
Cost of errors today223
Value total768
System access322
Rules and past cases321
Stability322
Feasibility total965
Cost of a wrong output231
Caught and reversible321
Decisions about people331
Risk total883
VerdictFirst roundNext roundRescope: automate the preparation, not the decision
ShapeFixed steps with an AI stepFixed branches, or a narrow agentFixed steps, with a person deciding

Supplier invoice coding would go in the first round. A mid-sized distributor receives invoices as PDFs by email and codes each one to a ledger account and cost centre. An AI step reads the invoice, a deterministic check matches it to the purchase order within a set tolerance, and anything outside tolerance goes to a person. Years of invoices already coded by hand make the test set, and a miscoded line can be reversed with a journal. Its value is only middling, but for the distributor the build connects to the ledger and sets up the logging and testing its later automations reuse.

Delivery-change requests are a next-round candidate. A retailer's customers ask by email and chat to move a delivery, change the address or cancel, and both the wording and the path vary: check the order, ask the carrier, rebook, perhaps refund a delivery fee. Feasibility falls short of the first-round bar, so this belongs after a first build that has connected the retailer's order system. Before anyone builds an agent, map a few hundred past requests. If nearly all of them follow a handful of paths, build those paths as fixed branches with an AI step that reads the request; if they do not, a narrow agent with three or four tools fits, with any refund above a set amount sent to a person.

Hardship payment plans fail as full automation, and that is the useful finding. A utility's customers in financial difficulty ask for a payment plan. The decision is significant for the person, the rules leave room for judgement and the customer may be vulnerable, so the candidate trips the stop check on decisions about people before any score is added, and the scores agree. Rescoped, it looks different: an AI step assembles the case, summarising the account history and the customer's own statement, and drafts the letter, while a trained member of staff makes the decision. That preparation work is a candidate in its own right and should be scored again as one.

When every case still needs a person

Some automations need a person to check every output before it goes anywhere. For high-risk work that is often right, and the fourth principle of the AI Playbook asks for humans to validate any high-risk decisions influenced by AI. The saving then depends on checking being much faster than doing the work, which is worth measuring before anyone counts it.

The failure I watch for is a review queue that is nearly always right. A reviewer who has approved ninety-nine cases in a row tends to read the hundredth less carefully, and the check turns into a signature. Two measures show whether that is happening: how often reviewers change an output, and how often a sample of approved cases turns out to be wrong.

The better design sends people the cases that need them. Put deterministic checks after the AI step, such as totals that must reconcile, fields that must match a record and limits that must not be exceeded, and send the failures plus a random sample to review. That is the pattern where the model proposes and a deterministic check decides, and it gives the reviewer a real job rather than a rubber stamp.

Public bodies carry one more duty. Central government departments, and the arm's length bodies within its scope, must publish records of the algorithmic tools they use in decision-making under the Algorithmic Transparency Recording Standard, and the AI Playbook encourages other public sector organisations to use it too.

A process scoring worksheet you can copy

Use one copy per candidate process. Fill in the evidence before the score, so every score rests on something another person can check.

Process scoring worksheet
Process scoring worksheet: [process name]

Process owner: [name, role, hours a week they can give]
Systems it reads from: [ ]
Systems it writes to: [ ]

1. Stop checks (any 'no' pauses the candidate)
Named owner with time: yes / no
The process is agreed and stable: yes / no
Past cases with known outcomes, or a rule that confirms an output: yes / no
No significant decision about a person made by software alone: yes / no

2. Value (score 1 to 3, with the evidence)
Volume a month: [ ] (source: [ ]) score: [ ]
Time per case: [ ] minutes (source: [ ]) score: [ ]
Cost of errors today: [ ] (source: [ ]) score: [ ]
Value total: [ ] / 9

3. Feasibility (score 1 to 3, with the evidence)
System access: [API / export / none] score: [ ]
Rules and past cases: [share of past cases the written rules decide correctly] score: [ ]
Stability: [planned changes in the next year] score: [ ]
Feasibility total: [ ] / 9

4. Risk (score 1 to 3, where 3 is safest)
Cost of a wrong output: [ ] score: [ ]
Caught and reversible: [the check that catches it, how it is undone] score: [ ]
Decisions about people: [ ] score: [ ]
Risk total: [ ] / 9

5. Shape
Can the people doing the work draw the path? yes / no
Shape: fixed steps / fixed steps with an AI step / agent
Steps that need a person: [ ]

6. Round: first / next / later, with checks in front / not yet
Reason: [ ]
Re-score on: [date]

Where 1AYM fits

If you have a long list and no agreed order, our AI Opportunity & Feasibility Sprint produces one. It is fixed scope and typically takes two to four weeks: we look at how the work runs today and what state the data and systems are in, score the candidates on value, feasibility and risk, and write the sequenced roadmap, the investment case and the governance conditions.

If the first process is already chosen, what follows is build work: the connections to your systems, the AI step, the checks around it and the logs that show what happened on every case. We take small fixed-scope statements of work as well as larger builds, a signed scope can start within a day, and if you already have a scoped job we can resource it on contract from the collective of associates who work with 1AYM, held to the same standard. Either way, the first step is a 30-minute call.

For engineers: measuring candidates, building the AI step and scoping an agent

The scores are only as good as the evidence behind them. These are the checks I would run before a first build and keep running during it.

Mine the event log
Pull case-level timestamps from the ticketing, ERP or workflow system to get volume, handling time and the paths cases actually take. The spread of distinct paths in the log is the evidence for fixed steps or an agent.
Replay past cases against the rules
Write the rules as code and run them over past cases with known outcomes. The share they decide correctly is the part that needs no model at all; the rest is what the AI step has to handle, and its size is your real scope.
Structured outputs, validated
Have the AI step return a fixed schema and validate it before anything downstream reads it: types, allowed values, totals that reconcile and IDs that exist in the system of record.
Idempotent writes and dry runs
Key every write on a stable ID from the source, so a retry after a timeout cannot post twice, and give every destructive action a dry-run mode that shows the change before it lands.
Scope agent tools tightly
OWASP's LLM06:2025 Excessive Agency traces the risk to excessive functionality, excessive permissions and excessive autonomy. Give an agent the fewest tools with the narrowest permissions, run each action in the requesting user's security context, and require a person to approve high-impact actions.
Evaluation set in CI
Keep the past cases and their right answers in the repository, and rerun them on every change to a prompt, a model or a rule, with thresholds that fail the build.
Shadow mode before cut-over
Run the automation beside the current process on live inputs, compare the outputs, and cut over one case type at a time. Keep a switch that sends everything back to people.
One trace per case
Record the inputs, model and prompt versions, tool calls, check results, the reviewer's decision and the final write for every case. The same record serves as the audit trail and the debugging tool.

Sources

  1. [1]UK government, Artificial Intelligence Playbook for the UK Government (10 February 2025), read 29 September 2026
  2. [2]Anthropic, Building effective agents (19 December 2024), read 29 September 2026
  3. [3]OpenAI, A practical guide to building agents (PDF, April 2025), read 29 September 2026
  4. [4]GOV.UK, Data (Use and Access) Act 2025: data protection and privacy changes (published 27 June 2025), read 29 September 2026
  5. [5]ICO, The Data (Use and Access) Act 2025: what does it mean for organisations? (updated 19 June 2026), read 29 September 2026
  6. [6]OWASP GenAI Security Project, LLM06:2025 Excessive Agency, read 29 September 2026
  7. [7]GOV.UK, Algorithmic Transparency Recording Standard hub (updated 8 May 2025), read 29 September 2026

Frequently asked questions

What is AI workflow automation?

Software that runs a business process across your systems, following fixed steps where the path is known and using an AI model for the parts that need reading or judgement, such as extracting fields from a document or classifying a request. Exceptions and consequential decisions go to a person.

Which processes should a business automate with AI first?

Ones that run often, take real time per case, draw on systems you can connect to, have past cases with known outcomes to test against, and where a wrong output is caught or can be undone. Invoice coding, request triage and document extraction often fit. Significant decisions about individuals should keep a person deciding.

Does a workflow need an AI agent?

Usually not at first. If the people doing the work can draw the path, build it as fixed steps with an AI step where reading or judgement is needed. An agent fits work whose steps change from case to case, and it costs more per case and is harder to test.

Should the first process be the most valuable one?

Not necessarily. The first build also sets up the connections, logging and tests that later automations reuse, so a process that is easy to connect and safe to get wrong often makes a better first project. Take on the high-value, high-risk processes once that groundwork has run for a full business cycle.

How is AI workflow automation different from RPA?

RPA bots drive application screens the way a person would, following recorded steps. AI workflow automation usually connects to systems through their APIs and adds a model where the input is unstructured. The two can run side by side, and the question for an existing bot is whether to keep it, extend it with an AI step or replace it.

Further

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM