Buyer guide
AI implementation: what the first 30 days should deliver
The first 30 days of an AI implementation should end with evidence you can open, run or read. Whoever does the work, whether an AI implementation consultant, a consultancy or your own team, by day 30 you should be able to see five things: working access to the agreed systems and data, one real workflow running end to end in your environment, an evaluation against real cases with the results written down, the first controls on access, approval and logging, and a handover pack already under way. If no use case has been chosen yet, the first month is a short discovery instead, and it should end in a ranked list of what to build and a costed plan.
First working slice. One real workflow taken end to end in the buyer's own systems and against its own data, small enough to build in weeks and complete enough to test, secure and hand over.
Checked . The public guidance this page cites was read on each publisher's own site on this date. Guidance changes, and the ICO says its DPIA guidance is under review.
What a good first month produces, week by week
Treat the weeks below as a shape to hold a supplier to. Nobody can promise them on your behalf, because access, data quality and approvals set the pace far more often than engineering does. That is where I see first months stall, and it is why a good plan starts all three on day one. Each week should end with something you can open, run or read, and a week that ends with only a status update is a week to ask about.
This guide assumes you have already chosen who will do the work. If you have not, run the four checks in our buyer guide to choosing an enterprise AI consultancy, linked at the end of this page, on any pitch first.
| Dimension | What you should see | What to ask for | Warning sign |
|---|---|---|---|
| Week 1: access and data | Accounts, environments and data access for one agreed workflow, working rather than requested. A baseline of how the work performs today: time, error rate or cost per case. | The written scope of the first slice, the measure it will be judged on, and the list of access requests with a named owner for each. | Access requests still unsent at the end of the week, or no baseline, so no later result can be compared with anything. |
| Week 2: a working first slice | One real workflow running end to end in your environment against your data, even if it is rough. The code and configuration in your repositories. | A live run you watch, on cases you choose, and a note of the riskiest assumption and whether the slice has tested it yet. | A demonstration on the supplier's laptop, in the supplier's accounts or on sample data. |
| Week 3: evaluation | A test set built from real cases your process owner has checked, with pass thresholds agreed before any results came in. | The results, failures included, and a way to rerun the same tests on every change. | A few good examples picked after the fact, or an accuracy figure with no test set behind it. |
| Week 4: controls and handover | Least-privilege access, a human approval step for high-risk actions, logs of what the system did, and a way to switch it off. | The runbook, the decision log, documentation of the data, models and prompts, and the named owner on your side for each part. | Security and governance scheduled for a later phase, or knowledge that exists only in the supplier's heads. |
| Day 30: the decision | A written recommendation to continue, change course or stop, with the evidence from weeks one to four attached. | The cost and plan for the next stage, and what it would take for your team to run the slice without the supplier. | A request for more time with no new evidence, or a scope that has grown before the first slice works. |
Week 1: access, data and a baseline
Much of week one is spent waiting on other people, so it has to start on the first morning. The supplier needs accounts, somewhere to run code and read access to the data for one workflow. Tell your security, data protection and legal contacts about the work now, while their questions are still cheap to answer, rather than at go-live. The UK government's AI Playbook, written for public servants but just as useful outside government, tells teams to engage compliance, legal and data protection experts early, and to work with commercial colleagues from the start.
If the workflow touches personal data, screen it for a data protection impact assessment (DPIA) this week. The ICO's guidance says a DPIA must be carried out before processing that is likely to result in a high risk to people, and recommends doing one if you are in any doubt. The ICO also says that guidance is under review following the Data (Use and Access) Act. This section describes the guidance as published on 29 September 2026. It is not legal advice.
Then measure the work as it runs today: the time, error rate or cost per case. The AI Playbook recommends establishing a baseline of metrics to measure project outcomes against. Skip it and the day-30 review turns into an argument about impressions that nobody can settle.
Week 2: one real workflow, end to end
The first slice is one workflow, run through your real systems against your real data. It is allowed to be rough. What it cannot do is skip the hard integration, because the integration is usually where the risk sits, and a demonstration that avoids it tells you very little about your project.
Pick the slice that tests the riskiest assumption first. The government's Service Manual asks teams building early prototypes to do the same: identify the riskiest assumptions and test those. If the system of record has no usable API, or the data is not good enough to act on, you want to learn that in week two rather than month three.
Ask for the code, prompts and configuration to live in your repositories and your cloud accounts from the first commit. Moving them later is slow and fiddly. A supplier who resists it is telling you something about how the contract will end.
Week 3: evaluation you can rerun
If an AI system has not been tested against real cases, it has not been tested. Build the test set from cases your process owner has already decided, awkward ones included, and agree what counts as a pass before anyone sees a score. Set the threshold after the results and it tends to land wherever the results did.
Ask for two kinds of number: how well the model performs, and whether the service meets users' needs and business goals. A model can score well while the people meant to use it ignore its output, so the second number is the one that tells you whether the project was worth paying for. The AI Playbook makes the same split between model metrics and service metrics, and warns that the two can be very different.
The tests should run again on every change to a prompt, a model or a rule, so the evaluation still works after the supplier has gone. The AI Playbook asks teams to test fully before deployment and to keep checking the live tool. The NCSC's guidelines for secure AI system development make the same point about updates: changes to data, models or prompts can change how a system behaves.
Week 4: controls and the handover pack
Controls belong in the first month. Adding them to a system that is already running costs more than building them in, and a later phase has a habit of never arriving. Ask for these four by name.
- Least privilege
- The system acts under its own identity, with only the access the workflow needs. The NCSC guidelines ask providers to apply least privilege and to restrict the actions an AI component can trigger in other systems.
- Human approval
- A named person approves high-risk actions before they happen. The AI Playbook asks that humans validate any high-risk decisions influenced by AI. Our write-up of the pattern where the model proposes and a deterministic check decides shows one way to build that step.
- Logs
- A record of inputs, outputs and approvals. The NCSC guidelines ask providers to log inputs to support audit, investigation and remediation, and to protect logs as sensitive data.
- A way to stop
- A switch that turns the capability off, a route for anyone to report a problem and an incident plan behind it. The NCSC guidelines expect incident response, escalation and remediation plans that cover AI systems.
Start the handover pack now. Written in the last week of a contract, it tends to record what people remember rather than what was decided. It should hold the runbook, a dated decision log and documentation of the data, models and prompts. The NCSC guidelines list what that documentation covers, including data sources, intended scope and limitations, guardrails and potential failure modes. The AI Playbook adds a plan for knowledge transfer and clear roles, including who has the authority to change the code.
Warning signs in the first month
Any one of these deserves a direct conversation with the supplier. Two or more in the same month is a reason to stop and re-plan before you spend any more.
- Nobody asked for access on day one
- Access, data and approvals take the longest to arrange. A plan that starts them in week three has already lost most of the month.
- The demonstration runs on their machine
- Sample data in a supplier's account shows a model can do the task on someone's data. It does not show that your workflow works.
- An accuracy figure with no test set
- A percentage with no list of cases behind it cannot be checked or rerun.
- Controls are left for phase two
- Access rules, approval steps and logs added after launch cost more and are easier to skip.
- Your code lives in their accounts
- If the repository, the prompts and the cloud resources are not yours from the first commit, ask why.
- The scope grows before anything works
- Adding use cases before the first one runs looks like progress and usually delays it.
- Organisation-wide change promised in 30 days
- A month buys one working slice and the evidence to decide on the next. Ask for the evidence behind any bigger claim.
What your organisation has to provide
No supplier, however good, can make up for a missing owner or missing access. Have these in place before day one, or at least named.
- A sponsor who can decide
- Someone with budget authority who will make the day-30 decision and can unblock access when it stalls.
- A process owner with time
- The person who does or runs the work today, available every week to explain it, supply cases and check results.
- Access, started early
- Accounts, environments, data and the approval route, with a named owner for each request.
- Real cases with known answers
- Past examples where the right outcome is already agreed. They become the test set in week three.
- Security, data protection and legal contacts
- Named people who review the work while it is being built, not only at the end.
- An owner for afterwards
- The team that will run the system once the supplier leaves. A handover needs someone to receive it.
A day-30 review checklist you can copy
Share this with the supplier at kick-off and fill it in together on day 30. Any line you cannot complete is the agenda for the review.
Day-30 review: [workflow name] Date: [date] Sponsor: [name, role] Process owner: [name, role] 1. Access. Every account, environment and data source the slice needs is working. Open requests: [list, each with an owner]. 2. Baseline. How the work performs today: [time, error rate or cost per case], measured on [date]. 3. Working slice. The workflow runs end to end in our environment against our data. Seen running on [date] on cases we chose. 4. Ownership. Code, prompts, test sets and infrastructure are in our repositories and accounts: [locations]. 5. Evaluation. Test set of [number] real cases. Pass threshold [value], agreed on [date]. Result: [score]. Failures listed in [location]. The tests rerun on every change: yes / no. 6. Controls. Own identity with least privilege: yes / no. Human approval for high-risk actions: yes / no. Logs of inputs, outputs and approvals: yes / no. Switch-off tested: yes / no. 7. Data protection. DPIA screening done on [date]. Outcome: [not required / completed / in progress]. 8. Handover. Runbook, decision log and documentation of data, models and prompts: [locations]. Our owner after handover: [name]. 9. Decision. Continue / change course / stop. Reason: [one paragraph]. Cost and plan for the next stage: [reference].
Where 1AYM fits
At 1AYM the first paid piece is usually a fixed-scope opportunity and feasibility sprint of two to four weeks, ending in a ranked list of what to build, the architecture for it and an investment case, or a fixed-scope first production slice against one real workflow. Work can start within a day of the scope being signed, and the access, data and approval requests go out on day one, because the delay on these projects is almost never the model. Expect weeks rather than quarters to a first working slice.
We work in three stages: discover, ship a working slice, then run it governed. Each ends with something written down and handed over, so you can stop after any stage and still own what you paid for. Where the contract needs it, copyright in the deliverables is assigned to you. For a longer programme, we work as a retained implementation team alongside your engineers, typically two to three days a week under one statement of work. Our about page covers how an engagement with 1AYM runs. If you want a second opinion on a supplier's first-month plan, whoever wrote it, book the call below.
For engineers: what the first slice should include
The four weeks as an engineering reviewer would check them. Every item below should be visible in the repository, the cloud account or the logs, so none of it rests on the supplier's word.
- Repositories and accounts
- Code, prompts, evaluation sets and infrastructure definitions live in the client's source control and cloud accounts from the first commit. Supplier access runs through the client's identity provider, so removing it is one step rather than a clean-up project.
- Environments
- A non-production environment with access to real data for the one workflow, segregated from other sensitive stores. No production writes until the week-four controls exist.
- Identity
- A dedicated service identity per AI component, scoped to the workflow. Where the source system supports it, use delegated user identity, so the system never holds more access than the person it acts for.
- Evaluation in CI
- A versioned test set of real cases, with expected outcomes and thresholds committed to the repository. Any change to a prompt, model, tool definition or rule that drops a score below its threshold fails the build.
- Gates on actions
- Deterministic checks on every proposed write: schema, business rules and reconciliation against the system of record. Failed checks and high-risk action classes route to a named approver.
- Logging
- One structured record per action: input reference, model and prompt version, check results, approver and timestamp. Keep it access-controlled, because the NCSC guidance treats logs as sensitive data.
- Switch-off and rollback
- A flag that disables the capability without a deploy, and dry-run modes so a change can be inspected before it lands. Test the flag in the first month: the day-30 checklist asks whether anyone has.
- Documentation
- Model and data cards, or an equivalent, covering data sources, scope, limitations, guardrails, retention and known failure modes, in line with the NCSC's documentation guidance.
Sources
- [1]UK government, Artificial Intelligence Playbook for the UK Government (10 February 2025), read 29 September 2026
- [2]NCSC, Guidelines for secure AI system development: secure design (27 November 2023), read 29 September 2026
- [3]NCSC, Guidelines for secure AI system development: secure development, read 29 September 2026
- [4]NCSC, Guidelines for secure AI system development: secure deployment, read 29 September 2026
- [5]NCSC, Guidelines for secure AI system development: secure operation and maintenance, read 29 September 2026
- [6]GOV.UK Service Manual, How the alpha phase works, read 29 September 2026
- [7]ICO, When do we need to do a DPIA?, read 29 September 2026
Frequently asked questions
What should an AI implementation consultant deliver in the first 30 days?
One real workflow working in your own systems, tested against real cases, with the first controls on access, approval and logging in place and a handover pack started. If no use case has been chosen, the first month is a short discovery that ends in a ranked list of what to build and a costed plan. Either way, day 30 should end with evidence and a decision rather than a status report.
Is 30 days long enough to reach production?
For one narrow workflow it can be, if access and approvals move quickly. More often, day 30 brings a working slice in a non-production environment and the evidence to decide whether to take it live. If a supplier promises organisation-wide change in a month, ask exactly what will be running on day 30.
Who should own the code and prompts?
You should, from the first commit. They belong in your repositories and cloud accounts, with supplier access granted through your identity provider. Check that the contract assigns copyright in what the supplier produces to you. The UK government's AI Playbook advises buyers to consider intellectual property rights and strategies to avoid vendor lock-in when they write their requirements.
What if the first slice fails its evaluation?
Then the month has done its job. A failed evaluation with a clear cause, such as data that is not good enough or a system with no usable API, is cheap next to finding the same problem after launch. The day-30 decision should say whether to change the approach, pick a different workflow or stop.
Do we need a DPIA before we start?
If the workflow processes personal data in a way likely to result in a high risk to people, UK GDPR requires a DPIA before the processing starts, and the ICO recommends doing one if you are in any doubt. Screen for it in week one. This is not legal advice: your data protection officer or legal adviser should make the call.
Further
- How to choose an enterprise AI consultancy in the UK · Four checks to run on any pitch, before you hire anyone for the first month.
- Fractional AI implementation team · A retained senior team that owns the architecture, builds with your engineers and hands the platform over.
- AI readiness assessment · 13 questions on ownership, data, systems and controls, scored on the page, before you hire anyone.
- Production AI systems · The fixed-scope architecture and build, once the first slice has earned a production system.
- AI governance framework · The policy and controls a governed rollout carries forward from the first slice.
We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.
Last reviewed · 1AYM