Guide
AI workflow automation: which processes to automate first
Automate first the process that scores well on value, feasibility and risk at once. Value means it runs often, takes real time per case and its errors cost money today. Feasibility means its inputs sit in systems you can connect to, its rules can mostly be written down, and past cases with known outcomes exist to test against. Low risk means a wrong output is caught or can be undone before it harms anyone. The most valuable process is often not the right first one, because the first build also puts the connections, logging and tests in place for everything after it. Use fixed steps wherever the path can be drawn in advance, with an AI step only where the input needs reading or judging, and keep agents for work whose steps change from case to case.
AI workflow automation. Running a business process across systems with software that follows fixed steps, uses an AI model for the parts that need reading or judgement, and sends exceptions to a person.
1AYM is an OpenAI Select Partner. Its founder holds personal Claude certifications. 1AYM is not an Anthropic partner. This guide cites both vendors' published guidance on agents and ranks neither. 1AYM also builds AI integration and automation, one of the routes this guide describes, so weigh our view with that in mind.
Checked . The vendor guidance and the UK government, ICO and OWASP sources this page cites were read on the publishers' own sites on 29 September 2026. The worked examples are fictional: their processes and scores show the method and are not client results. Where the page touches data protection law it summarises published guidance and is not legal advice.
Three shapes of AI workflow automation
Ask three suppliers to quote for AI workflow automation and you can get three different builds under the same name. The difference decides the cost, the risk and how the thing gets tested, so settle which shape a process would take before you score it.
- Fixed steps
- Software follows a path drawn in advance: an invoice arrives, it is matched to the order, posted and sent for approval. No model is involved. Where the inputs are already structured this is often the whole answer, and it is the cheapest shape to run and to test.
- Fixed steps with an AI step
- The same drawn path, with a model inside the one or two boxes that need reading or judgement, such as pulling fields from a PDF, classifying an email or drafting a reply. The path, the permissions and the checks around the model stay fixed.
- An agent
- A model decides which steps to take, in what order and with which tools, until it judges the task done. It earns its place when the path differs from case to case, and it is harder to test and costs more per case.
This guide is about choosing and sequencing the processes. The build itself, with audit logs, dry runs and rollback on every write to your systems, is what our AI integration and automation services cover.
Who builds it is a separate choice. An AI automation agency usually builds on a no-code platform and often runs the result for a monthly fee, while a consultancy starts from the business problem. We set out the difference between an AI agency and an AI consultancy, including what you own when the contract ends, in a separate guide.
Score each candidate on value, feasibility and risk
Start the long list with the people who do the work, not with a vendor's list of use cases. The UK government's AI Playbook says the choice of use case must be led by business and user needs, pain points and inefficiencies, not by what the technology can do. That holds well outside government.
Then score each candidate from 1 to 3 on the nine criteria below, using measured figures wherever you have them. Risk is scored the other way round from the other two, so that 3 is always the better score: a 3 on risk means a wrong output is cheap, caught early and easy to undo.
| Dimension | What to look at | Scores 3 when | Scores 1 when |
|---|---|---|---|
| Value: volume | Cases a month, counted from the system that logs them rather than estimated | Hundreds or thousands a month, every month | A handful a month, or one burst a year |
| Value: time per case | Handling time, from system timestamps where they exist | Tens of minutes of skilled time per case | A minute or two |
| Value: cost of errors today | Rework, write-offs, penalties and complaints the current process causes | Errors today cost real money or customers | Errors are rare and cheap to put right |
| Feasibility: system access | Where the inputs arrive and where the outputs must be written | Every system involved has a usable API or a supported export | The work lives in email threads, personal spreadsheets or a screen with no API |
| Feasibility: rules and past cases | How much of the decision can be written down, and whether past cases with known outcomes exist to test against | The rules are written or can be, and hundreds of past cases have recorded outcomes | Nobody can say how the decision is made, and outcomes are not recorded |
| Feasibility: stability | How often the process, its forms or its systems change | Stable for at least the next year | Being redesigned, or a system it depends on is about to be replaced |
| Risk: cost of a wrong output | What a wrong answer does if it gets through | An internal record someone corrects | Money leaves, a customer is told something false, or a person is refused something |
| Risk: caught and reversible | Whether a check catches the error before it causes harm, and whether the action can be undone | A rule or a reconciliation catches it, and the action can be reversed | Nothing checks it, and the action cannot be undone, such as a payment or a sent letter |
| Risk: decisions about people | Whether the output decides something significant about an individual | No individual is affected | It decides someone's access to money, work, services or care |
Keep the three groups as three totals rather than adding them into one. A process scoring 9 on value and 3 on risk needs a different conversation from one scoring 6 on everything, and a single number hides the difference.
The scores rank candidates against each other. They are not a business case, and before a build is approved the chosen process still needs a measured baseline and a proper estimate of the return.
Four findings that stop a candidate before scoring
Some findings end the conversation for now, whatever the scores say. I check for these first, because any one of them can sink a project that looks excellent on paper.
- No named owner
- Nobody who runs the process today has the time to explain it, supply past cases and check the results. Without that person, the build guesses at the rules.
- A process nobody agrees on
- If the process is being redesigned, or the people in it disagree about how it should run, automating it makes a bad process faster. Settle the process first.
- No way to check an output
- No past cases with known right answers, and no rule that can confirm an output is correct. Build that test set before anything else; it is cheaper than finding the errors in production.
- Software alone deciding about a person
- Under UK GDPR as amended by the Data (Use and Access) Act 2025, a significant decision based solely on automated processing needs safeguards: telling people about the decision, letting them make representations and challenge it, and letting them obtain human intervention. The ICO notes that the wider choice of lawful bases for these decisions does not extend to special category data. Keep a person making the decision, or take advice before designing it any other way. This is not legal advice.
What goes first, what goes next
The process with the highest value score is often the wrong first choice. A first build has two jobs: to return something, and to lay the groundwork every later automation reuses, meaning the connections to your systems, the logging, the habit of testing against past cases and the routes to a reviewer. A process that is easy to connect and safe to get wrong does the second job well, even when its value is only middling.
So I would sequence the list in rounds, re-scoring the rest after each one, because the first build changes what is feasible for everything after it.
- First
- Feasibility and risk totals of 7 or more, and value of at least 6: something that runs every day, where a wrong output is caught before it matters. The aim is a working system in production with its logs and tests, not a showcase.
- Next
- Candidates that reuse the systems the first build connected, including ones with messier inputs or more risk. Take on more risk only once the first automation has run for a full business cycle and its error rate has been measured.
- Later, with checks in front
- High value and high risk, such as anything that pays out money or decides something for a customer. These go last, with a deterministic check between the AI step and the write, and a person on every consequential case.
- Not yet
- Anything with a feasibility total of 5 or less. Improve the data, the system access or the process itself, and score it again in a quarter.
When a workflow needs an agent, and when fixed steps are enough
Anthropic's engineering guidance draws the line clearly. In its terms, workflows are systems where models and tools are orchestrated through predefined code paths, and agents are systems where the model directs its own process and tool use. It recommends finding the simplest solution possible and adding complexity only when needed, and it notes that agents trade latency and cost for better task performance, with the potential for compounding errors.
OpenAI's practical guide to building agents reaches the same place from the other side. It suggests prioritising workflows that have resisted automation: decisions full of judgement and exceptions, rule sets that have become costly to maintain, and heavy reliance on unstructured data such as documents and conversation. If a use case does not clearly meet those criteria, it says, a deterministic solution may suffice.
My own test is simpler. If the people doing the work can draw the flowchart, build the flowchart and put a model in the boxes that need reading or judgement. Reach for an agent when the number and order of steps change from case to case, and a drawn path would need dozens of branches to cover what a person does. In the first year of an automation programme I would expect many more of the first kind than the second.
| Dimension | Fixed steps with an AI step | An agent |
|---|---|---|
| The path through the work | Drawn in advance; the model fills in one step | Chosen by the model while it runs |
| Testing | Each AI step is tested on its own against past cases, and the path is tested like any other software | Whole runs have to be tested, because the path itself can go wrong |
| Cost per case | Predictable: a known number of model calls | Varies with how many steps the model takes |
| How it fails | One step gives a wrong output, which a check after it can catch | A wrong early choice can carry through every later step |
| Permissions | Each step gets only the access that step needs | The agent's tools set the limit, so each tool needs the narrowest access possible |
| Typical fits | Invoice coding, email triage, document extraction, drafting replies for a person to send | Requests and investigations whose steps differ each time, within a small set of tools |
Three candidates, scored (fictional)
These three processes, from three different organisations, are invented to show the method. The organisations, volumes and scores are illustrative, and none of them is a client result or a benchmark. Each is scored from 1 to 3 on the nine criteria above, with risk scored so that 3 is safest.
| Dimension | Supplier invoice coding | Delivery-change requests | Hardship payment plans |
|---|---|---|---|
| Volume | 3 | 3 | 2 |
| Time per case | 2 | 1 | 3 |
| Cost of errors today | 2 | 2 | 3 |
| Value total | 7 | 6 | 8 |
| System access | 3 | 2 | 2 |
| Rules and past cases | 3 | 2 | 1 |
| Stability | 3 | 2 | 2 |
| Feasibility total | 9 | 6 | 5 |
| Cost of a wrong output | 2 | 3 | 1 |
| Caught and reversible | 3 | 2 | 1 |
| Decisions about people | 3 | 3 | 1 |
| Risk total | 8 | 8 | 3 |
| Verdict | First round | Next round | Rescope: automate the preparation, not the decision |
| Shape | Fixed steps with an AI step | Fixed branches, or a narrow agent | Fixed steps, with a person deciding |
Supplier invoice coding would go in the first round. A mid-sized distributor receives invoices as PDFs by email and codes each one to a ledger account and cost centre. An AI step reads the invoice, a deterministic check matches it to the purchase order within a set tolerance, and anything outside tolerance goes to a person. Years of invoices already coded by hand make the test set, and a miscoded line can be reversed with a journal. Its value is only middling, but for the distributor the build connects to the ledger and sets up the logging and testing its later automations reuse.
Delivery-change requests are a next-round candidate. A retailer's customers ask by email and chat to move a delivery, change the address or cancel, and both the wording and the path vary: check the order, ask the carrier, rebook, perhaps refund a delivery fee. Feasibility falls short of the first-round bar, so this belongs after a first build that has connected the retailer's order system. Before anyone builds an agent, map a few hundred past requests. If nearly all of them follow a handful of paths, build those paths as fixed branches with an AI step that reads the request; if they do not, a narrow agent with three or four tools fits, with any refund above a set amount sent to a person.
Hardship payment plans fail as full automation, and that is the useful finding. A utility's customers in financial difficulty ask for a payment plan. The decision is significant for the person, the rules leave room for judgement and the customer may be vulnerable, so the candidate trips the stop check on decisions about people before any score is added, and the scores agree. Rescoped, it looks different: an AI step assembles the case, summarising the account history and the customer's own statement, and drafts the letter, while a trained member of staff makes the decision. That preparation work is a candidate in its own right and should be scored again as one.
When every case still needs a person
Some automations need a person to check every output before it goes anywhere. For high-risk work that is often right, and the fourth principle of the AI Playbook asks for humans to validate any high-risk decisions influenced by AI. The saving then depends on checking being much faster than doing the work, which is worth measuring before anyone counts it.
The failure I watch for is a review queue that is nearly always right. A reviewer who has approved ninety-nine cases in a row tends to read the hundredth less carefully, and the check turns into a signature. Two measures show whether that is happening: how often reviewers change an output, and how often a sample of approved cases turns out to be wrong.
The better design sends people the cases that need them. Put deterministic checks after the AI step, such as totals that must reconcile, fields that must match a record and limits that must not be exceeded, and send the failures plus a random sample to review. That is the pattern where the model proposes and a deterministic check decides, and it gives the reviewer a real job rather than a rubber stamp.
Public bodies carry one more duty. Central government departments, and the arm's length bodies within its scope, must publish records of the algorithmic tools they use in decision-making under the Algorithmic Transparency Recording Standard, and the AI Playbook encourages other public sector organisations to use it too.
A process scoring worksheet you can copy
Use one copy per candidate process. Fill in the evidence before the score, so every score rests on something another person can check.
Process scoring worksheet: [process name] Process owner: [name, role, hours a week they can give] Systems it reads from: [ ] Systems it writes to: [ ] 1. Stop checks (any 'no' pauses the candidate) Named owner with time: yes / no The process is agreed and stable: yes / no Past cases with known outcomes, or a rule that confirms an output: yes / no No significant decision about a person made by software alone: yes / no 2. Value (score 1 to 3, with the evidence) Volume a month: [ ] (source: [ ]) score: [ ] Time per case: [ ] minutes (source: [ ]) score: [ ] Cost of errors today: [ ] (source: [ ]) score: [ ] Value total: [ ] / 9 3. Feasibility (score 1 to 3, with the evidence) System access: [API / export / none] score: [ ] Rules and past cases: [share of past cases the written rules decide correctly] score: [ ] Stability: [planned changes in the next year] score: [ ] Feasibility total: [ ] / 9 4. Risk (score 1 to 3, where 3 is safest) Cost of a wrong output: [ ] score: [ ] Caught and reversible: [the check that catches it, how it is undone] score: [ ] Decisions about people: [ ] score: [ ] Risk total: [ ] / 9 5. Shape Can the people doing the work draw the path? yes / no Shape: fixed steps / fixed steps with an AI step / agent Steps that need a person: [ ] 6. Round: first / next / later, with checks in front / not yet Reason: [ ] Re-score on: [date]
Where 1AYM fits
If you have a long list and no agreed order, our AI Opportunity & Feasibility Sprint produces one. It is fixed scope and typically takes two to four weeks: we look at how the work runs today and what state the data and systems are in, score the candidates on value, feasibility and risk, and write the sequenced roadmap, the investment case and the governance conditions.
If the first process is already chosen, what follows is build work: the connections to your systems, the AI step, the checks around it and the logs that show what happened on every case. We take small fixed-scope statements of work as well as larger builds, a signed scope can start within a day, and if you already have a scoped job we can resource it on contract from the collective of associates who work with 1AYM, held to the same standard. Either way, the first step is a 30-minute call.
For engineers: measuring candidates, building the AI step and scoping an agent
The scores are only as good as the evidence behind them. These are the checks I would run before a first build and keep running during it.
- Mine the event log
- Pull case-level timestamps from the ticketing, ERP or workflow system to get volume, handling time and the paths cases actually take. The spread of distinct paths in the log is the evidence for fixed steps or an agent.
- Replay past cases against the rules
- Write the rules as code and run them over past cases with known outcomes. The share they decide correctly is the part that needs no model at all; the rest is what the AI step has to handle, and its size is your real scope.
- Structured outputs, validated
- Have the AI step return a fixed schema and validate it before anything downstream reads it: types, allowed values, totals that reconcile and IDs that exist in the system of record.
- Idempotent writes and dry runs
- Key every write on a stable ID from the source, so a retry after a timeout cannot post twice, and give every destructive action a dry-run mode that shows the change before it lands.
- Scope agent tools tightly
- OWASP's LLM06:2025 Excessive Agency traces the risk to excessive functionality, excessive permissions and excessive autonomy. Give an agent the fewest tools with the narrowest permissions, run each action in the requesting user's security context, and require a person to approve high-impact actions.
- Evaluation set in CI
- Keep the past cases and their right answers in the repository, and rerun them on every change to a prompt, a model or a rule, with thresholds that fail the build.
- Shadow mode before cut-over
- Run the automation beside the current process on live inputs, compare the outputs, and cut over one case type at a time. Keep a switch that sends everything back to people.
- One trace per case
- Record the inputs, model and prompt versions, tool calls, check results, the reviewer's decision and the final write for every case. The same record serves as the audit trail and the debugging tool.
Sources
- [1]UK government, Artificial Intelligence Playbook for the UK Government (10 February 2025), read 29 September 2026
- [2]Anthropic, Building effective agents (19 December 2024), read 29 September 2026
- [3]OpenAI, A practical guide to building agents (PDF, April 2025), read 29 September 2026
- [4]GOV.UK, Data (Use and Access) Act 2025: data protection and privacy changes (published 27 June 2025), read 29 September 2026
- [5]ICO, The Data (Use and Access) Act 2025: what does it mean for organisations? (updated 19 June 2026), read 29 September 2026
- [6]OWASP GenAI Security Project, LLM06:2025 Excessive Agency, read 29 September 2026
- [7]GOV.UK, Algorithmic Transparency Recording Standard hub (updated 8 May 2025), read 29 September 2026
Frequently asked questions
What is AI workflow automation?
Software that runs a business process across your systems, following fixed steps where the path is known and using an AI model for the parts that need reading or judgement, such as extracting fields from a document or classifying a request. Exceptions and consequential decisions go to a person.
Which processes should a business automate with AI first?
Ones that run often, take real time per case, draw on systems you can connect to, have past cases with known outcomes to test against, and where a wrong output is caught or can be undone. Invoice coding, request triage and document extraction often fit. Significant decisions about individuals should keep a person deciding.
Does a workflow need an AI agent?
Usually not at first. If the people doing the work can draw the path, build it as fixed steps with an AI step where reading or judgement is needed. An agent fits work whose steps change from case to case, and it costs more per case and is harder to test.
Should the first process be the most valuable one?
Not necessarily. The first build also sets up the connections, logging and tests that later automations reuse, so a process that is easy to connect and safe to get wrong often makes a better first project. Take on the high-value, high-risk processes once that groundwork has run for a full business cycle.
How is AI workflow automation different from RPA?
RPA bots drive application screens the way a person would, following recorded steps. AI workflow automation usually connects to systems through their APIs and adds a model where the input is unstructured. The two can run side by side, and the question for an existing bot is whether to keep it, extend it with an AI step or replace it.
Further
- AI integration and automation service · The build once the process is chosen: AI and automation inside the systems you run, with audit logs, dry runs and rollback.
- AI strategy and roadmap · The fixed-scope sprint that scores candidate use cases on value, feasibility and risk and sets the order.
- AI agents for business · When a process does call for an agent: where agents fit, where they are the wrong tool and the controls to set first.
- AI ROI: estimate it, then check it · Turn the value score into a business case: baseline, formulas, the costs people forget and a stress test.
- RPA vs AI automation · When to keep an RPA bot, extend it with an AI step or replace it.
- When Zapier is not enough · The signs a no-code workflow has outgrown its platform, and the options.
- Agentic workflow design · The checks that decide whether an automated write is safe to make.
We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.
Last reviewed · 1AYM