# 1AYM full text > 1AYM is an enterprise AI and platform consultancy and an OpenAI Select Partner. We take organisations from AI ambition to governed production systems their people use every day. Generated from the same content modules that render https://www.1aym.com, so nothing here can describe a page that does not exist. British English throughout. Last reviewed 1 September 2026. Clients are not named in this file. Where a figure came from our own measurement rather than a client publication, the source line says so, exactly as it does on the page. ## Key Facts - Type: enterprise AI and platform consultancy, headquartered in Stoke-on-Trent, United Kingdom. Language: en-GB. - Delivery: UK, the Gulf and the US. Registered in England and Wales. - Company status: OpenAI Select Partner, awarded after a partner agreement, compliance review and technical assessment. 1AYM is not an Anthropic or Claude partner, and "OpenAI certified" is not a thing 1AYM is, or claims. - Certifications: Claude Certified Architect – Professional (Anthropic, 23 July 2026) and Claude Certified Associate – Foundations (Anthropic, 21 July 2026), held by the founder and Principal Architect, independently verifiable on Credly. - Compliance: we work to UK GDPR and the Data Protection Act 2018. We hold no ISO or SOC certification. - Clients: referred to by descriptor in this file and on every route, with two exceptions. One engagement file names its client, because the published source it links carries the name. The homepage logo bar shows client marks. - Capability areas: 9. Named engagements: 5. Published engagement files: 5. Client environments in daily production use: 2. ## Who we are The work is executive discovery, architecture, hands-on engineering, governance, rollout and measurement. Where we work. We are based in Stoke-on-Trent, West Midlands, and deliver across the UK, the Gulf and the US. One client environment runs on Postgres in Google Cloud’s Doha region (me-central1) to meet Gulf data-residency requirements, and another reached 91% adoption in North America, a figure published in the Anthropic customer case study. Who does the work. Every 1AYM engineer is either certified on the platforms we build on or has shipped inside a top-tier engineering organisation: Meta, Spotify, UBS, Starling Bank, S&P Global, Sky. We scale the team to the contract, and the people who write the strategy write the code. The senior engineers a client meets first: a Principal Architect (10+ years; AI adoption, agent systems, Claude and OpenAI platform architecture; holds both Anthropic Claude certifications; in production at the client environments below); a Principal Engineer (14 years, London; data platform, context and memory architecture; shipped inside Meta, Spotify, Viasat/Inmarsat); a Principal Data Engineer (17 years, London; pipeline orchestration, CI/CD, data infrastructure; shipped inside Starling Bank, UBS, S&P Global, Sky, Société Générale). The Principal Architect is Tayyeb Mahmud, who founded the practice. 1AYM delivers the work rather than any one person; other colleagues are named only with their permission, and clients meet the engineers before signing. Meta, Spotify, Viasat/Inmarsat, Starling Bank, UBS, S&P Global, Sky and Société Générale are former employers of our engineers, not clients of 1AYM. Certifications and status are two different things and 1AYM keeps them apart. The Claude Certified Architect – Professional (Anthropic, issued 23 July 2026) and Claude Certified Associate – Foundations (Anthropic, issued 21 July 2026) are held by 1AYM’s founder and Principal Architect, and are independently verifiable on Credly, which carries the holder’s public record. 1AYM’s OpenAI Select Partner status is a company status, held by the practice rather than by any individual, and was awarded after a partner agreement, compliance review and technical assessment. "OpenAI certified" is not a thing 1AYM is, or claims: the correct wording is "OpenAI Select Partner and Claude certified". 1AYM is not an Anthropic or Claude partner. No vendor bias. Model choice at 1AYM comes from production evidence across OpenAI, Anthropic, Gemini and open-source, not from a vendor deal. We are model-selective across OpenAI, Anthropic, Gemini and open-source components, chosen per workload rather than by vendor loyalty. In production today: OpenAI audio models for speech assessment with a Gemini fallback, Claude for enterprise platform and agent estates, Codex and Claude Code for engineering work. We work to UK GDPR and the Data Protection Act 2018, and hold no ISO or SOC certification. ## Homepage (https://www.1aym.com/) **Enterprise AI your people use every day.** The people who write the strategy write the code. Every 1AYM engineer is certified or has shipped inside Silicon Valley, banking and government-scale systems, and we scale the team to the contract. OpenAI Select Partner and Claude certified. We run OpenAI, Anthropic and Gemini models in production, and open-source where it fits, so the choice for your use case comes from evidence rather than a vendor deal. We are based in Stoke-on-Trent, West Midlands, and deliver across the UK, the Gulf and the US. ### How it works Three stages, in this order. Each one ends with something written down and handed over, so you can stop after any of them and still own what you paid for. #### Discover Typically two to four weeks with your sponsor and the people who do the job today. We agree what the outcome is worth, what stands in the way, and how it will be judged. You leave with a ranked list of what to build, the architecture for it, and an investment case with real numbers in it. #### Ship a working slice One real workflow, end to end, in your systems and against your data. A demonstration that skips the hard integration proves nothing, so we do not build one. You leave with working software in your environment, and evidence of what it does. #### Run it, governed Access rules, approval gates, monitoring and training, then a staged rollout. Your team runs it. We measure adoption, cycle time and cost, so the change is reported rather than asserted. You leave with a running system, a team that can maintain it, and numbers for the board. ### What you can buy Five named engagements. Each one has a buyer, a stated output or a clear commercial shape, and paperwork you can put in front of procurement without a rewrite. #### AI Opportunity & Feasibility Sprint For C-suite and transformation leads. Fixed scope · typically 2–4 weeks. Find out which two or three workflows are worth building, what each is worth and what it would cost, before you commit a budget to any of them. You get: - A ranked shortlist of use cases, with the value and the risk against each one - The architecture, the data and integration constraints, and which models fit - A delivery roadmap and an investment case your finance team can sign off #### Enterprise AI Platform & Agentic Workflows For CIO, CTO, CDO, operations and finance. Fixed-scope architecture, then a production build. Turn scattered pilots into one governed place for AI to run, then put real workflows on it. Access rules, checks and human approval are built in, so the second workflow does not start from scratch the way the first one did. You get: - Platform architecture: who can see what, and the context an AI system is allowed to use - Agents that act on real systems, checked against fixed rules and approved by a human before anything lands - Testing, monitoring and an operating model your team owns #### AI Engineering Transformation (Codex & Claude Code) For engineering leaders. Typically 6–12 weeks · enablement, guardrails and CI. Your engineers are already using AI tools, with or without a policy. This puts review, guardrails and CI around that, so the code it writes reaches production the same way every other change does. You get: - Codex and Claude Code workflows mapped to how your teams actually ship - Repository guardrails, CI gates and multi-agent delivery patterns - Training, documentation and adoption figures you can report on #### Fractional AI Platform Architect For scale-ups and enterprises. Retained · typically 2–3 days a week · one statement of work. Senior architecture and delivery ownership without hiring a permanent AI platform leader. We hold the technical decisions, and the standard the work is judged against. You get: - Architecture ownership, with the model and vendor decisions written down and dated - Delivery governance and technical review of work already in flight - Team enablement, so the capability stays when the engagement ends #### Embedded Engineers on Contract For programmes that need senior capacity now. Contract · from three months · priced per day. A senior engineer, or a team, inside your programme on contract. You direct the work day to day, and we hold the standard every 1AYM engineer is hired against, plus the cover behind them. You get: - A named senior engineer or a team inside your programme, typically within weeks - The 1AYM standard, and cover across the stream so no role hangs off one person - A clean handover, or conversion to a fixed-scope build once the work is scoped ### Production proof Client environments depend on this work every working day. Every figure is either the client's published case-study number or our own count or measurement on the systems we run, and each one says which. #### A PE-backed marketing agency We lead the architecture of the enterprise AI platform underneath the agency's organisation-wide AI programme. That platform is featured in a published Anthropic customer case study, which describes the platform. What we built: - The publishing pipeline and release controls that let non-technical staff author AI skills and Claude run them - An OAuth-secured organisational-context MCP server with default-deny permissions and a full audit trail - Automated access provisioning from the HR system, and reporting on what AI actually costs - A re-architecture of the skills estate to the Anthropic SDK's structure, with reference files loaded on demand and MCP servers that inject targeted context in place of whole files, which by our own measurement cut the average cost per session by 60% - ~1,000 · Employees reached (source: Anthropic customer case study) - ~400 · Skills authored in four weeks (source: Anthropic customer case study) - 91% · Adoption in North America (source: Anthropic customer case study) #### A government-accredited EdTech We hold end-to-end technical ownership of the live production estate of a Middle East education company: the exam platform, the APIs and the data, replatformed onto Postgres 16 in Google Cloud's Doha region to meet Gulf data-residency requirements. Its IELTS speaking assessment is not one model call wrapped in a product. It is a measuring instrument, built to Ministry of Education evidence standards and in daily use by the client's own team. What we built: - Whole-interview speaking assessment: the official IELTS band descriptors applied in ordinary code, so every band can be reproduced and inspected line by line - Five layers of evidence, from deterministic acoustic measurement to a phoneme-level pronunciation model, with the AI never doing the arithmetic - Confidence routing to a human examiner, and accent fairness and agreement with examiners as release gates the build has to pass before it ships - 3,100+ · Students on the platform (source: Production database, 21 August 2026) - 1,000+ · Added in the last 90 days (source: Production database, 21 August 2026) - 1,400+ · Automated tests, from ~300 (source: Our count, in CI) #### Measured outcomes, client unnamed - £2–4M · Forecast annual saving (source: Our model of token consumption, reviewed by the client's finance team) - 46% · Faster first response (source: Our measurement · client unnamed) - In the meeting · Governed finance answer, from ~2 hours (source: Our measurement · client unnamed) - 82% · Less manual triage (source: Our measurement · client unnamed) ### 1AYM vs the alternatives There are four ways to buy this kind of work. They differ in who writes the code, how long it takes to reach production, and what you still own once the invoice is paid. | | 1AYM | Big-4 / systems integrator | Staff augmentation | Strategy-only advisory | | --- | --- | --- | --- | --- | | Who does the work | Certified engineers, or engineers who shipped inside Meta, Spotify, UBS and Starling Bank. Senior, hands-on, no bench of juniors. | A partner sells it; a large mixed-seniority team delivers it. | People you brief, manage and direct yourself. | Advisers who hand the build to somebody else. | | Time to first production slice | Typically weeks. One real workflow, end to end, in your environment. | Months, after a discovery phase and a mobilisation phase. | As fast as the plan you give them, and no faster. | None. The output is a document. | | Model choice | From production evidence across OpenAI, Anthropic, Gemini and open-source, not from a vendor deal. | Usually shaped by the vendor alliance and reseller economics the firm already has. | Whatever the people you hired happen to know. | A recommendation, with nothing in production behind it. | | Governance | Built as platform controls: default-deny access, permission-aware context, deterministic checks, human approval, evaluation gates in CI. | A governance workstream, usually a policy set and a steering committee. | Yours to define, enforce and audit. | A framework for somebody else to implement. | | What you are left with | Running systems, the architecture written down, and a team that can extend both. | A system your team did not build, and a change-request route into it. | Whatever got built. The knowledge leaves when they do. | A roadmap, and the build still to buy. | | Commercial shape | Fixed-scope statements of work, a retained architecture engagement, or embedded engineers from three months. One supplier, one contract. | Multi-phase programme against a rate card. | Priced per person, per day, open-ended. | Fixed fee for the report. | ### Six things we do not trade None of them is a methodology. They are the constraints we hold on every engagement, and the reason the work is still running after we leave. - **Certified or proven**: Every engineer is certified on the platforms we build on, or has shipped inside a top-tier engineering organisation. - **Fixed scope**: Defined outputs, a defined price and acceptance criteria in the statement of work, or a clear contract. - **Governed from day one**: Access control, approval gates and evaluation in CI are part of the build, not a later project. - **Your people keep the keys**: Architecture, code and operating model are documented and handed over. Nothing depends on us staying. - **No vendor bias**: Chosen per workload across OpenAI, Anthropic, Gemini and open-source, on production evidence rather than vendor loyalty or a vendor deal. - **Overlapping coverage**: The engineers on your work cover each other's streams, so one person's calendar is never the critical path. ### Certified, or proven at the top Every engineer here is either certified on the platforms we build on, or has shipped inside a top-tier engineering organisation: Meta, Spotify, UBS, Starling Bank, S&P Global, Sky. We scale the team to the contract. - **Tayyeb Mahmud · Founder and Principal Architect** (10+ years) · AI adoption, agent systems, and Claude and OpenAI platform architecture. Sits with the sponsor, sets the standard, and still writes the code. In production: A PE-backed marketing agency · a government-accredited EdTech. Holds both Anthropic Claude certifications, Architect – Professional and Associate – Foundations, verifiable on Credly. - **Principal Engineer** (14 years · London) · Data platform work, and the context and memory architecture that decides what an AI system is allowed to see, keep and recall. Shipped inside: Meta · Spotify · Viasat/Inmarsat. - **Principal Data Engineer** (17 years · London) · Pipeline orchestration, CI/CD and data infrastructure: the layer most AI programmes are actually blocked on once the model choice is settled. Shipped inside: Starling Bank · UBS · S&P Global · Sky · Société Générale. We name colleagues only with their permission, and we do not name clients. 1AYM delivers the work, and the role, the years and the organisations on each card are the standard being met. You meet the engineers before you sign anything. 1AYM is an enterprise AI and platform consultancy and an OpenAI Select Partner, founded by Tayyeb Mahmud. Based in Stoke-on-Trent, West Midlands, we deliver across the UK, the Gulf and the US. We take organisations from AI ambition to governed production systems their people use every day. Client environments depend on that work today. ### Frequently asked questions #### Who actually does the work? Certified engineers, or engineers who have shipped inside Meta, Spotify, UBS, Starling Bank, S&P Global and Sky, and the same people who scoped the work. There are no juniors on it. Our principal architect leads the architecture and holds the Claude Certified Architect – Professional and Claude Certified Associate – Foundations certifications, both verifiable on Credly. A principal engineer of fourteen years takes the data platform and the context and memory layer; a principal data engineer of seventeen takes pipeline orchestration, CI/CD and data infrastructure. They cover each other's streams, and we scale the team to the contract. 1AYM delivers the work rather than any one person, we name colleagues only with their permission, and you meet the engineers before you sign anything. #### How do engagements start? With a 30-minute call to work out whether your problem is one we should take. If it is, the first paid piece is fixed-scope: usually a two-to-four week opportunity and feasibility sprint, or a first production slice against one real workflow. We will say plainly when the answer is that you do not need us. Expect weeks rather than quarters to a first working slice, because the delay on these projects is almost never the model. It is access, data quality and the approval path, so we start those on day one. #### How do you contract? Three shapes. Most work is a fixed-scope statement of work with defined outputs and a defined price: a sprint, an architecture and build, or an enablement programme. The second is a retained architecture engagement, typically two to three days a week, under a single statement of work. The third is embedded engineers, a senior engineer or a team inside your own programme on a contract from three months, where the constraint is capacity rather than a scoped deliverable. In every case it is one supplier and one contract. A statement of work carries the outputs, the price and the acceptance criteria; an embedded contract carries the roles, the minimum term and the review cadence. Where a programme needs more engineers we scale the team to the contract, and every one of them is held to the same standard. #### What does “production” mean to you? A system real people depend on in their working day. That means identity and access control, monitoring, deterministic checks where the output matters, a human approval path for anything high-risk, and a named owner inside your organisation. Something that runs on an engineer's laptop is a prototype. Client environments in production today meet that definition. #### How do you handle governance and data? Governance is built as controls rather than written as a policy: least-privilege and default-deny access, permission-aware context so an AI system sees only what the user is already entitled to see, audit trails that record refusals as well as answers, deterministic validation, human approval sized to risk, and evaluation gates in CI. We work to UK GDPR and the Data Protection Act 2018, plus whatever your sector adds. Where data must stay in a region we architect for it: one production estate runs on Postgres in Google Cloud's Doha region. We hold no ISO or SOC certification. #### Which models do you use? Whichever one the workload argues for. The choice comes from what we run in production across OpenAI, Anthropic, Gemini and open-source, not from a vendor deal. We are model-selective across OpenAI, Anthropic, Gemini and open-source components, chosen per workload rather than by vendor loyalty. In production today: OpenAI audio models for speech assessment with a Gemini fallback, Claude for enterprise platform and agent estates, Codex and Claude Code for engineering work. Where a model is a poor fit for a job we will say so and price the alternative. ## Services (https://www.1aym.com/capabilities) Index of the nine capability areas. The nav calls them Services; the route is /capabilities. ### AI Opportunity & Feasibility Sprint (https://www.1aym.com/capabilities/executive-ai-discovery) Reference: C-01. Question this page answers: How do we turn board-level AI ambition into a buildable plan? Sold as: AI Opportunity & Feasibility Sprint. Fixed scope · typically 2–4 weeks. Executive AI discovery turns an ambition into a plan a CFO can fund and an engineer can start on Monday. We run the leadership sessions, map where the value sits, test feasibility against your data and systems, settle build versus buy, and write the roadmap, the investment case and the governance conditions. Programmes rarely stall because the technology is hard. They stall because nobody turned the ambition into something specific, sequenced and fundable, and the cheapest place to find that out is before the budget is committed. #### The problem this solves A board asks for an AI strategy. What comes back is either a slide deck with no engineering in it, or an engineering plan with no business case in it. Both stall, for the same reason: the person writing them could only see one half of the problem. Discovery closes that gap. It produces a plan that a CFO can fund and an engineer can start on Monday, because the same person wrote both halves and can defend the trade-offs between them. #### How the work runs It is deliberately short. The aim is a decision, not a documentation exercise. - **Leadership sessions**: What the business is actually trying to change, commercially rather than technically. Where the pressure is coming from, and what success would look like to the people funding it. - **Workflow analysis**: How the work is done today, including the spreadsheets and manual steps nobody documents. This is usually where the real opportunities are hiding. - **Data and systems feasibility**: What data exists, what state it is in, who is allowed to see it, and which systems would have to be touched. Most AI plans die here, and it is cheaper to find out early. - **Opportunity map**: Candidate use cases scored on value, feasibility and risk, so the sequence is defensible rather than whichever idea had the loudest sponsor. - **Governance review**: What has to be true for legal, security, and risk to sign off, established before anything is built rather than discovered at the end. #### Why it comes first Almost every expensive AI failure traces back to skipping this. A team builds the use case that was easiest to describe, rather than the one that was most valuable, and finds out at rollout that the data was not accessible or the process owner was never consulted. Discovery is the cheapest step in the programme and the one that determines whether the rest of the money is well spent. #### What you get - Opportunity map with value, feasibility and risk scoring - Current-state workflow analysis - Technical architecture for the recommended direction - Sequenced delivery roadmap - Risk and governance notes - Build-versus-buy recommendation #### Start here if - Leadership has committed to AI but nobody has converted it into a plan - Several teams are running disconnected experiments - A business case is needed before budget is released - Previous AI work stalled and no one is certain why #### Questions ##### How long does discovery take? It is scoped as a short, focused engagement rather than an open-ended consulting phase: long enough to interview the people who own the work and assess the data honestly, short enough that it does not become the project. The output is a decision and a roadmap, not a research programme. ##### Do you need access to our data to do this? It helps a lot. Feasibility claims made without seeing the data are guesses. Where access is impossible in the timeframe, the assessment states its assumptions so they can be tested before build. Last reviewed 29 August 2026. ### AI harness & platform engineering (https://www.1aym.com/capabilities/ai-harness-platform-engineering) Reference: C-02. Question this page answers: How do we enable our engineers on Codex, Claude Code and Cursor? Sold as: AI Engineering Transformation (Codex & Claude Code). Enablement, guardrails and CI · typically 6–12 weeks. Vendors build the harness. 1AYM sells architecture, skills packages, guardrails and CI on Codex, Claude Code and Cursor, so AI-assisted changes take the same path to production as every other change, and so the company can move work between those tools without rewriting the operating layer. We do not build a proprietary harness in place of Claude, ChatGPT or Cursor. Those vendors maintain the runtime; we make it usable, governed and portable across the estate. The engagement is still sold as AI Engineering Transformation. #### What the engagement covers The client already has vendor harnesses. The missing piece is the layer that makes them safe to depend on across teams: how work is structured, which skills are shared, what the model may touch, and how a change is reviewed and shipped. - **Architecture**: Which harness is used for which job, how the company's systems sit relative to it, and what is shared rather than rewritten per team. - **Skills packages**: Reusable skills the vendor harness can load, so a workflow is owned once and used on Codex, Claude Code or Cursor as the work requires. - **Guardrails**: Repository rules, permissions and the actions that need a person, enforced in the path to production rather than as a policy slide. - **CI and review**: The same gates every other change already passes, so AI-assisted work is not a side channel. - **Enablement**: Training and the operating model your engineers keep when the engagement ends. #### Built on the tools already in use, not instead of them Engineers are already in Codex, Claude Code or Cursor, with or without a policy. The engagement starts there. It maps those workflows to how the teams actually ship, then adds the architecture and the gates, rather than asking anyone to abandon the vendor harness for a house-built substitute. That sequencing keeps the work honest: every control exists because a real repository and a real release process demanded it. #### What you get - Architecture for how vendor harnesses are used across the estate - Skills packages reusable on Codex, Claude Code and Cursor - Repository guardrails and permission model - CI gates and the review path to production - Enablement, documentation and an operating model your team owns #### Start here if - Engineers are already using Codex, Claude Code or Cursor, and nothing consistent sits around that - Each team has invented its own skills, permissions and review rules - AI-assisted changes reach production by a different path from everything else - A second vendor harness is arriving and the first one's setup cannot move with it #### Questions ##### Does this lock us into one model provider? The opposite, if the architecture is written to sit on the harness rather than inside one vendor's product. Codex, Claude Code and Cursor remain the runtimes. Skills packages and guardrails are what you take with you when the mix changes. ##### Can we keep the tools our engineers already use? That is the starting point. The engagement maps Codex, Claude Code and Cursor to how the teams ship, then adds architecture, skills and CI around them. A replacement harness is not the offer. Last reviewed 29 August 2026. ### Production AI systems (https://www.1aym.com/capabilities/production-ai-systems) Reference: C-03. Question this page answers: How do we get an AI system into production and keep it there? Sold as: Enterprise AI Platform & Agentic Workflows. Fixed-scope architecture plus production build. A demonstration has one happy path, a clean input and somebody watching. A production system has every input your business can generate, nobody watching, and a consequence attached to being wrong. We build the second kind: internal copilots, workflow agents, voice agents and the middleware underneath them, designed around who is allowed to see what, what happens when a model fails, what gets logged, and where a human signs off. That is the part that decides whether it is still running in a year. #### The gap between a demo and a system A demo has one happy path, a curated input, and a human watching. A production system has every input the business can generate, no one watching, and a consequence attached to being wrong. Closing that gap is most of the engineering. It means deciding what happens when the model is unavailable, when the input is malformed, when the output fails validation, when a downstream system rejects a write, and when someone needs to know six months later why a particular decision was made. #### What the design actually covers - **Data access and permissions**: What the system can read and write, acting on behalf of whom, and how that is enforced rather than assumed. - **Failure modes**: What happens on timeout, malformed output, rate limit, partial write, and downstream rejection, each one designed rather than discovered. - **Evaluation**: Regression suites that run when prompts or models change, so a quality drop is caught before users find it. - **Human review**: Where a person sits in the loop, what they are shown, and how their corrections feed back. - **Auditability**: A record of what ran, on which inputs, producing which output: the thing that makes the system answerable afterwards. #### Safe by default The engineering defaults matter more than the model choice: idempotent operations so a retry cannot double-write, dry-run modes so a change can be inspected before it lands, audit logs as a by-product rather than an afterthought, and kill switches that work. These are unglamorous and they are the difference between automation that is clever and automation that is safe to depend on. #### What this looks like in production Three figures from systems we run, each one published in full on the engagement file it links to. - ~1,170 automated tests cover the speaking pipeline on one production estate, among them a regression suite where disabling any one guardrail breaks specific frozen cases. (source: Our count, in CI; https://www.1aym.com/work/education-ai-enablement) - ~300 to 1,400+ the automated test estate on that same platform, over the course of the engagement. (source: Our count, in CI; https://www.1aym.com/work/education-ai-enablement) - 60% lower average cost per session on another client's AI platform after we re-architected its skills estate. (source: Our measurement on the platform; https://www.1aym.com/work/ai-platform-enablement) #### What you get - Working system running against production data - Evaluation and regression suites - Human review paths and escalation routes - Audit logging and observability - Runbook and handover documentation #### Start here if - A prototype works but nobody will let it near production - An AI feature shipped and quietly degraded - There is no way to tell whether output quality has changed - The system needs to write to a system of record #### Questions ##### How do you evaluate an AI system that has no single right answer? By separating the parts that do from the parts that do not. Structure, schema, and factual grounding can be checked deterministically. Genuinely subjective quality is scored against a fixed set of graded examples, so the question becomes whether today's output is worse than last week's rather than whether it is perfect. ##### What happens when the model provider changes the model underneath us? This is exactly what the regression suite exists for. Without one, a silent provider-side change is discovered by users. With one, it is caught on the next run and the routing layer can pin or reroute while the difference is assessed. Last reviewed 22 August 2026. ### Agentic workflow design (https://www.1aym.com/capabilities/agentic-workflow-design) Reference: C-04. Question this page answers: How do we make AI automation dependable enough for finance or compliance? A model that is nearly always right sounds excellent until it runs against your ledger all day. Nearly right is wrong, repeatedly, and delivered with exactly the same confidence as the correct answers. So the model is never the last thing to touch a decision. The agent proposes; deterministic rules, schemas, tests and a human where the risk warrants it decide what proceeds. Finance, data migration and compliance work is built this way because an auditor, not a user, finds the errors. #### Why accuracy is the wrong measure A model that is right 95% of the time produces fifty wrong entries a day at a thousand runs, delivered with the same confidence as the correct ones. The useful question is not how often the system is right, but whether anything notices when it is not. The pattern that answers it is set out in full on the concept page, with worked detail on what a verifier actually checks. - **The agent proposes**: The judgment across messy inputs. - **The verifier gates**: Deterministic checks, none of them a model. - **Pass proceeds, fail holds**: A failure routes to a person with the reason attached. #### Where it earns its keep Finance-critical automation, data migration validation, knowledge extraction, semantic-layer query validation, report generation, operational triage, and internal tooling that writes to systems of record. It is unnecessary where the output is read by the person who asked for it before anything happens: a drafting assistant does not need a gate, because the reader is one. #### What you get - Workflow design with explicit gate criteria - Deterministic validator suite - Human review routing and exception handling - Audit trail generated as a by-product of the gates - Test coverage for the verification layer #### Start here if - Automation touches money, identity, or a regulated process - A previous automation shipped errors nobody caught - Reviewers are rubber-stamping a queue that is mostly correct - The process must be explainable to audit or risk #### Questions ##### What do you need from us before gate design can start? The workflow, and the cost of a silent error in it. Those two facts decide how much of the check can be deterministic and where a person has to stand. Everything else follows from them. ##### Does this work on a process we have already automated? Usually, and it is the more common starting point. An automation that already runs has a history of the errors it produced, which is the best specification anyone could write for what the verifier has to catch. Last reviewed 22 August 2026. ### Data platform & AI enablement (https://www.1aym.com/capabilities/data-platform-ai-enablement) Reference: C-05. Question this page answers: How do we connect AI to data the business actually trusts? The expensive failure in enterprise AI is a plausible answer that disagrees with the official number by a few percent, for reasons nobody can trace. Two of those and the finance team stops trusting the tool, whatever the model behind it. We build the layer that prevents it: warehouse integrations, definitions agreed once in a semantic layer, data contracts and reconciliation, so an AI answer lands on the same figures as the board pack. The trust is built in the data layer, underneath the model. #### The trust problem The most common failure in enterprise AI is not a wrong answer. It is a plausible answer that disagrees with the official number by four percent, for a reason nobody can trace. Once that happens twice, the system is dead. People stop using it, because checking its output costs more than doing the work manually. Trust is the actual product, and it is built at the data layer rather than the model layer. #### What the work involves - **Warehouse integration**: Making Snowflake, BigQuery or the equivalent reachable by AI systems under the same access controls that govern everyone else. - **Semantic-layer enablement**: Exposing the business's own definitions (what counts as an active customer, a booked deal, a period), so a model inherits them rather than inventing its own. - **Data contracts**: Explicit agreements about shape, meaning and freshness, so an upstream schema change surfaces as a failed contract rather than a silently wrong answer. - **Reconciliation workflows**: Checks that an AI-produced figure agrees with the system of record, run automatically rather than trusted. - **AI-accessible reporting**: Reporting surfaces designed to be queried by a system, not only rendered for a human. #### Why semantics matter more than retrieval Most retrieval problems in enterprise AI are really definition problems. The model finds the right table and still produces the wrong number, because the business has three definitions of revenue and the model picked one. Fixing that at the semantic layer fixes it for every use case at once. Fixing it in prompts fixes it in one place until someone writes a new prompt. #### What you get - Warehouse and semantic-layer integration - Data contracts covering shape, meaning and freshness - Reconciliation checks against systems of record - AI-accessible reporting surfaces - Documentation of definitions the AI layer inherits #### Start here if - AI answers disagree with the official numbers - The business has multiple definitions of the same metric - An upstream schema change silently broke a downstream answer - Analysts are re-checking everything the AI produces #### Questions ##### Do we need a semantic layer before we can use AI on our data? Not before, but you will end up building one. Without shared definitions, each use case encodes its own interpretation in prompts and queries, and they drift apart. Formalising the definitions once is cheaper than reconciling five accidental versions of them later. ##### Is this the same as building a RAG system? No. Retrieval decides which documents or rows the model sees. This decides what the numbers in them mean and whether they agree with the system of record. A retrieval system on top of undefined data retrieves the wrong number faster. Last reviewed 27 July 2026. ### Enterprise integrations & automation (https://www.1aym.com/capabilities/enterprise-integrations-automation) Reference: C-06. Question this page answers: How do we put AI inside the tools our teams already use? AI in a separate tab does not change how work gets done. Nobody switches to a chat window for a task their existing system already half handles, so the value appears when it shows up in the ticket, the ledger or the record. We build it into the systems you already run: NetSuite, Xero, Slack, Notion, identity and HR feeds. Then we make it safe to depend on: dry-run modes, idempotent writes, rollback paths, audit logs, and never more access than any person in the company holds. #### Integration is the product AI that lives in a separate tab does not change how work gets done. People will not context-switch into a chat window to do a task their existing system already half-handles. The value appears when the capability shows up inside the workflow: in the ticket, the ledger, the channel, the record. That makes integration the deliverable rather than the plumbing behind it. #### The unglamorous parts that decide whether it survives - **Idempotency**: A retry must not double-write. Systems fail midway more often than they fail cleanly, and a non-idempotent sync turns a transient error into a data problem. - **Dry-run modes**: Every destructive operation should be inspectable before it lands, so a change can be reviewed rather than trusted. - **Rollback paths**: A defined way back from a bad run, decided before the bad run rather than during it. - **Audit logs**: What changed, when, on whose behalf, and why: a by-product of running rather than a compliance project afterwards. - **Rate limits and backoff**: Enterprise APIs throttle. Handling that properly is the difference between a sync that completes and one that half-completes nightly. #### Identity and permissions Anything that reads or writes on a user's behalf inherits that user's permissions, which means identity is part of the integration rather than a layer above it. SCIM provisioning, role mapping, and de-provisioning all have to work, including the unhappy paths. Getting this wrong is how an automation ends up with more access than any human in the organisation. #### What you get - Integration pipelines with idempotent writes - Dry-run and rollback tooling - Audit logging and run observability - Identity and permission mapping, including de-provisioning - Scheduling with predictable, monitored runs #### Start here if - AI capability exists but nobody uses it because it is in another tool - A sync fails partway and leaves inconsistent state - No audit trail for automated writes - Access granted to an automation that nobody reviews #### Questions ##### Can you work with our existing integration platform? Usually yes. Where a tool like a workflow automation platform already handles the orchestration adequately, the sensible move is to use it and add the missing guarantees around it. Custom middleware is worth building when the guarantees matter more than the convenience: typically idempotency, audit, and permission handling that low-code tools do not express well. ##### How do you handle systems with no usable API? By being honest about the trade-off. Scheduled exports, file drops, and database-level integration are all legitimate when an API is absent, provided the same guarantees hold: idempotency, dry runs, audit. Screen-scraping a system of record is normally where we would advise against automating at all. Last reviewed 27 July 2026. ### Fractional AI platform architect (https://www.1aym.com/capabilities/fractional-ai-platform-architect) Reference: C-07. Question this page answers: How do we get senior AI architecture ownership without hiring a permanent platform leader? Sold as: Fractional AI Platform Architect. Retained · typically 2–3 days a week · one statement of work. A permanent AI platform leader takes two quarters to find, sign and land, and a programme without one makes its architecture decisions by accident in the meantime. A fractional engagement puts that ownership in place now. We hold the target architecture, the vendor and model decisions, the technical review of work already in flight and the delivery governance, and we enable your own engineers to take all of it over. It is retained rather than fixed-scope, typically two to three days a week under one statement of work, and every decision is written down and dated so the reasoning outlives the engagement. #### The problem this solves A programme with a budget, a mandate and nobody who owns the architecture. The decisions still get made. They get made by whoever is in the room: the loudest vendor, the team with the nearest deadline, or the last proof of concept that happened to work. A permanent hire fixes it eventually. Two quarters of search, notice and ramp is the usual cost, and most of the decisions that shape the platform are taken before that person arrives. #### How the engagement runs A standing cadence, so the ownership is real rather than an advisory line on a slide. - **Architecture ownership**: We hold the target architecture and the trade-offs behind it, and we are answerable for them the way an employed lead would be. - **A weekly decision log**: Each decision, the options weighed and the reason for the one taken, written down and dated. It is the artefact that outlasts the retainer. - **Review of work in flight**: Technical review of what your teams and your suppliers are already building, so a problem is caught while it is still a design question. - **Vendor and model decisions**: Which model, which platform, and where the build-versus-buy line falls, with the production evidence behind each choice stated rather than asserted. - **Enablement**: The point of the engagement is that it ends. Your engineers take the architecture, the log and the operating model, and run them without us. #### What you are left with An architecture and a decision record your own team can defend to a board, to an auditor or to the next supplier, and engineers who helped write both. A written handover, so the end of the retainer is a date in the contract rather than a cliff. #### When it is the wrong choice When the problem is already scoped, a fixed-scope engagement is cheaper and faster. A feasibility sprint settles what to build; an architecture and production build ships it. Both are priced against a defined output, which a retainer is not. It is also the wrong shape when what is missing is capacity rather than ownership. That is a contract for embedded engineers, and it has its own page below. #### What you get - Target architecture, owned and kept current - A dated decision log covering vendor, model and build-versus-buy choices - Technical review of work already in flight - Delivery governance: gates, review points and release criteria - Enablement so your own engineers take the architecture over - A written handover at the end of the retainer #### Start here if - A funded AI programme with no single owner of the architecture - Vendor recommendations are the only technical opinion in the room - A permanent platform lead has been open for a quarter or more - Work is in flight and nobody senior is reviewing the design #### Questions ##### How much of the week does a retained architect hold? Typically two to three days a week, under one statement of work. The shape matters more than the total: ownership that only appears at an escalation is advice rather than ownership, so the days are a standing cadence. ##### Who makes the final call on a model or a vendor? You do. We hold the decision, set out the options and the reasoning, and put a date on it. What that buys is a record you can defend a year later, and an argument settled on evidence from production rather than on who presented last. ##### What stops this becoming a permanent dependency? The enablement sits inside the engagement rather than after it, and the handover is written as we go. Your engineers hold the architecture and the decision log alongside us, so the end of the retainer is a date rather than a risk to manage. Last reviewed 29 August 2026. ### Embedded Engineers on Contract (https://www.1aym.com/capabilities/embedded-engineers-on-contract) Reference: C-08. Question this page answers: How do we add senior engineers to a programme that is short of capacity rather than short of a plan? Sold as: Embedded Engineers on Contract. From three months · priced per day. We place a team, or a single senior engineer, inside your programme on a contract basis. Contracts run from three months and are priced per day. Every engineer meets the same standard as the rest of the practice: certified on the platforms we build on, or shipped inside a top-tier engineering organisation. You direct the work day to day, in your backlog and your stand-ups. 1AYM holds the standard and the cover behind it, so a stream never hangs off one person. The usual roles are an AI or platform architect, an AI engineer, and a data or platform engineer. #### The problem this solves A funded programme waiting on a hiring round. The plan is signed off and the budget is committed, and the start date now depends on how long it takes to find people. Or a delivery team short one senior role, with the date unmoved. Or a backlog that needs engineers rather than another roadmap. None of those is a defined output, so buying one as a fixed-scope engagement fits badly. #### How it runs One contract, a minimum of three months, on day rate commercials rather than against a deliverable. - **Who turns up**: A single senior engineer or a small team, inside your programme. The usual roles are an AI or platform architect, an AI engineer, and a data or platform engineer. - **The standard**: Every engineer meets the same bar as the rest of the practice: certified on the platforms we build on, or shipped inside a top-tier engineering organisation. - **Who directs the work**: You do, day to day, in your own backlog and your own stand-ups. We hold the standard and the cover, so a stream never hangs off one person. - **Weekly review**: One of our architects reviews the work each week, so an embedded engineer is never the only senior pair of eyes on a design. - **Swap or scale**: Changing a role, or adding one, happens under the same contract rather than through a new procurement round. #### What you are left with The work, in your repositories and your environments, built the way your own engineers build. Nothing depends on a system only we can reach. A clean handover at the end, or a conversion to a fixed-scope build once the work is defined enough to price against an output. #### When a fixed-scope engagement is the better buy When the outcome is defined, buy the outcome. A feasibility sprint settles what to build, and an architecture and production build ships it. Both are priced against a deliverable rather than a day, which is the cheaper trade whenever the deliverable can be written down. When what is missing is the technical ownership rather than the hands, a retained architect is the closer fit. #### What you get - A named senior engineer, or a team, inside your programme - One contract, from three months, priced per day - Weekly technical review by a 1AYM architect - Work delivered in your repositories, environments and process - Cover across the stream, so no role is a single point of failure - A clean handover, or conversion to a fixed-scope build #### Start here if - A funded programme is waiting on a hiring round - A delivery team is short one senior role and the date has not moved - The backlog needs engineers rather than another roadmap - Capacity is the constraint, and the outcome is not yet scoped enough to price #### Questions ##### How quickly can somebody start? Typically weeks rather than a hiring round. We match the roles to the programme first, because a fast start in the wrong discipline costs more than a slower one in the right discipline, and we say plainly when we do not have the right person for a role. ##### What happens if an engineer is not the right fit? We swap them, under the same contract. That is the practical difference between a contract with a practice and a direct hire: the cover sits with us, so a stream is never left waiting on one person's notice period or one person's holiday. ##### Can an embedded engagement become a scoped build? Often, and it is usually the better buy once it can be. Three months inside the programme is enough to know what the output is worth and what it takes, which is exactly what a fixed-scope statement of work needs before anyone can price it honestly. Last reviewed 22 August 2026. ### AI governance implementation (https://www.1aym.com/capabilities/ai-governance-implementation) Reference: C-09. Question this page answers: How do we actually implement AI governance rather than write a policy about it? Sold as: Enterprise AI Platform & Agentic Workflows. Fixed-scope architecture plus production build. Governance written as a policy document changes nothing about what a system does at three in the morning. We implement it as controls: identity and permissions the system enforces rather than assumes, deterministic verifier gates on consequential actions, evaluation running in CI, an audit trail produced as a by-product of those gates, human approval on the decisions that warrant it, and data residency architected in where a region is a requirement. It is built during the platform build, because retrofitting identity, gates and logging onto a running system costs more than designing them in. That is what makes the obligations UK GDPR and the Data Protection Act 2018 place on you testable instead of asserted. It is not a certificate: we hold no ISO or SOC certification, and controls are evidence, not a substitute for one. #### Why a policy document governs nothing An AI policy is a statement of intent: who is accountable, what is allowed, what has to be reviewed. None of it reaches the running system, because a system does not read policies. It reads permissions and configuration. The gap between the two is where the incidents live. A policy that forbids an agent from writing to a system of record is a sentence. A permission that stops it is a control. Governance an auditor can test is the second kind, and it is engineering work, not drafting work. #### What the policy needs underneath it Each one either exists in your estate or it does not, and anyone with access can check which. That is the difference between a control and an intention. Not every estate needs all six: a read-only assistant over public documents needs identity and logging, and the gates and the residency work start earning their cost when the system can write, or when the data is personal. - **Identity and permissions**: Every agent, service and person acts as somebody, with entitlements inherited from the systems that already hold them rather than a second permission model standing beside the first. - **Verifier gates on consequential actions**: Deterministic checks decide whether a proposed action is allowed to land. The pattern is set out in full on the concept page below; this is the engagement that puts it into your estate. - **Evaluation in CI**: Quality and safety checks run on every change to a prompt, a model or a rule, so a regression is caught by a build rather than by a customer. - **An audit trail**: The gates write the record as they run: what ran, on which inputs, under whose authority, and what it produced. It is a by-product of the check, so it exists whether or not anyone thought to ask for it. - **Human approval where it matters**: Approval sits on the decisions that warrant it, and the reviewer sees the specific claim they are being asked to accept. A queue of mostly correct items teaches people to click through it. - **Data residency**: Where a region is a requirement, the data stays in it by architecture, and we can say which region each store sits in. #### What this does for UK GDPR and the DPA 2018, and what it does not We work to UK GDPR and the Data Protection Act 2018, plus whatever your sector adds. The controls above are how those obligations become testable: lawful access is a permission, minimisation is a schema, a record of processing is an audit trail, and a decision somebody can contest is one whose evidence was kept. What they do not do is confer a certificate. Controls support your compliance position. They do not stand in for your lawful basis, your own assessments or your data protection officer's judgement, and no supplier can take those off you. We hold no ISO or SOC certification, and we do not imply one. If your security review makes one a hard requirement, we are the wrong supplier for you, and a first call is a cheaper place to find that out than the end of a scoring round. #### Where these controls have been built Governance is easy to assert and awkward to evidence, so here are four engagements where the controls were the deliverable rather than the appendix. - **Clinical research**: A part-built clinical-research platform taken to launch on a development environment with review gates and a release discipline its own team runs. Its documentation is aligned to MHRA, FDA and EU Annex 11 expectations, and no inspection or audit outcome is claimed for it or for us. - **Education assessment**: An IELTS speaking assessment where accent fairness and agreement with human examiners are release gates in CI, low confidence routes a case to a human examiner, and the rollout is staged behind controls. - **Finance permissions**: A connector answering finance questions against a governed semantic model while carrying each person's existing entitlements, so nobody gained access they did not already have. - **Identity provisioning**: Access to two tools driven from the HR record across a large membership matrix, with deterministic identity resolution, dry runs, an audit trail and a defined path back from a bad run. #### What you are left with An estate whose controls you can demonstrate on request, and governance documentation written from what the system does rather than from what somebody hoped it would do. Your engineers hold the gates, the evaluation suite and the audit trail, and can change all three. A control your supplier has to be in the building to operate is not a control you own. #### Controls already in the estates we run Each figure is published in full, with its own caveats, on the engagement file it links to. - 600+ groups, and more than 1,000 users, provisioned from the HR record, so access to Notion and Claude ends when somebody's record changes rather than when a colleague remembers. (source: Our delivery record; https://www.1aym.com/work/identity-provisioning-at-scale) - ~15% of interviews route to a human examiner on a confidence gate, with a further 5% audited at random whether the system was confident or not. (source: Our routing thresholds, as built; https://www.1aym.com/work/education-ai-enablement) - 7 regulator-facing governance documents drafted from the controls on that same estate, with the rollout staged behind gates rather than switched on. (source: Our delivery record; https://www.1aym.com/work/education-ai-enablement) - Unchanged the permission model on a finance connector. It carries each person's own Looker entitlements, so answering finance questions in the meeting widened nobody's access. (source: Our delivery record; https://www.1aym.com/work/finance-data-access-connector) #### What you get - A control map: every consequential action, its gate and its approver - Identity and permission model, enforced by the system rather than described - Deterministic verifier gates on the actions that can cause harm - Evaluation and regression suites running in CI on every change - Audit logging produced by the gates, with a stated retention position - Data residency recorded per store, and governance documentation written from the controls as built #### Start here if - An AI policy exists and nobody can show what enforces it - An agent can write to a system of record and nothing stops it - Risk, procurement or a regulator has asked how the system is controlled - Personal or clinical data is in scope and residency is a live question #### Questions ##### Do these controls make us compliant with UK GDPR? They support compliance rather than confer it. Lawful basis, retention, your record of processing and your own assessments stay yours. What changes is that the obligations become testable: a permission either exists or it does not, and an audit trail either holds the action or it does not, and neither answer depends on anyone's recollection. ##### Can you write the governance documentation as well as build the controls? Yes, and in that order, which is the order that makes it true. A document drafted first describes intentions. One written from an implemented estate describes behaviour, so the answer to what happens when an agent is asked to write to the ledger is the same in the document and in the code. ##### Does governance have to be a separate project? No, and it costs more when it is. Bought as its own workstream it arrives after the system and gets fixed to the outside of it. These controls are implemented inside the platform build, where identity, gates, evaluation and logging are design decisions rather than retrofits. ##### What if our security review asks for a certification you do not hold? Then the answer is no, and you have it before anyone spends a procurement cycle on it. We work to UK GDPR and the Data Protection Act 2018 and hold no ISO or SOC certification, and we do not imply one. What we put in front of a reviewer is the control set, the evidence each control produces and the engineers who built them. Last reviewed 1 September 2026. ## Engagement files (https://www.1aym.com/work) Written at a public-safe level: capability, constraints and numbers. Clients are not named; internal project names and proprietary business logic are held back. ### Self-serve finance answers, without breaking permissions (https://www.1aym.com/work/finance-data-access-connector) Reference: D-06. Client: Client withheld. Delivered by 1AYM. A question about client or financial performance used to take around two hours to come back from the finance team. It is now answered in the meeting where it is asked, and nobody gained access they did not already have: the connector carries each person's existing Looker entitlements, so people see exactly what they were already entitled to see. Connecting the data took an afternoon. The week that followed went on agreeing, with the finance director, the definitions that make an answer correct, and that week was the actual work. - ~2 hrs → in the meeting · Time to a governed finance answer - 1 afternoon · To connect the data - 1 week · To make the numbers trustworthy - Unchanged · Existing permission model #### The problem Financial information sat behind the finance team. Anyone elsewhere in the business who needed a figure, for a client conversation or a meeting or a decision, raised a request and waited, typically a couple of hours. The finance team spent a meaningful share of its week answering questions that were, in principle, already answered by data the company held. The obvious fix, letting people query the data directly, ran straight into the reason the gate existed: financial data is not uniformly shareable. Some of it is open to the business, some of it is privileged, and any solution that flattened that distinction was worse than the queue. #### What was built A connector between the company's AI workspace and its existing Looker setup, so questions asked in natural language resolved against the same governed model the business already used for reporting. - **Permission inheritance**: The connector carries the user's own Looker entitlements. Someone with privileged access keeps it; someone without sees only what was already open to them. No parallel permission model was created, because a second model is a second thing to get wrong. - **Grounded in the existing model**: Answers resolve against the company's canonical metrics rather than against raw tables, so the figures match the ones finance would have given. - **Semantic layer**: The definitions that make an answer correct (what counts as a client, a period, a booked figure), expressed once so every question inherits them. #### The hard part was not the integration Connecting the data took a single afternoon. That speed was itself the problem: it confirmed a belief held across the business that this kind of work is plug-and-play. It is not. A connector that returns a number is trivial. A connector that returns the *same* number the finance director would have given you is a week of work, and that week is the entire value. We spent it with the finance director, rapidly iterating the semantic layer and checking the agent's answers against his, until the two agreed on definitions the business actually uses. Had the project stopped after the afternoon, it would have shipped something that looked finished and quietly disagreed with finance, which is worse than the queue it replaced, because the queue was at least right. #### Outcome People pull the figures they need in the meeting where the question comes up, rather than preparing a request and waiting on someone else's queue. The finance team stopped being a lookup service for questions the data could answer on its own. The wider result was a corrected assumption. The engagement demonstrated to leadership that the gap between a working connection and a trustworthy one is where the actual engineering lives, which changed how subsequent AI work at the business was scoped. Stack: Looker, Semantic layer, AI workspace connector, Permission inheritance. Last reviewed 27 July 2026. ### An IELTS speaking assessment built as a measuring instrument (https://www.1aym.com/work/education-ai-enablement) Reference: D-02. Client: A government-accredited EdTech in the Middle East. Delivered by 1AYM. We hold end-to-end technical ownership of a government-accredited EdTech's live production estate, and its IELTS speaking assessment is the centre of it. Most AI speaking scorers are one model call wrapped in a product; this one is a measuring instrument. It listens to the whole interview, eleven to fourteen minutes of it, gathers acoustic and linguistic evidence independently, and applies the official band descriptors in ordinary code, so every band can be reproduced and inspected line by line. Where it is not confident it says so and routes the case to a human examiner. The pipeline is complete and independently verified, and accent fairness and agreement with human examiners are release gates whose validation is scheduled rather than results we are claiming. - 11–14 min · The whole interview is scored, not a sample - 5 layers · Of evidence, and the AI never does the arithmetic - ×3 · Examiner model runs, median taken, escalate on disagreement - ≥ 0.70 · Agreement with human examiners: a release gate, not a result - ~1,170 · Automated tests on the speaking pipeline #### Why this is the hard problem Two trained human examiners marking the same IELTS speaking interview agree at a correlation of about 0.90 [1]. That is the ceiling. An automated marker is measured against a standard that people do not hit perfectly themselves, so the useful question is not whether it is right every time. It is whether you can see how it reached a band, and catch it when it is wrong. The regulatory floor moved while this was being built. ETS's 2026 update to the TOEFL marks all eleven speaking items on the test as AI scored [2]. Ofqual fined Cambridge English £875,000 over automated-marking errors that ran undetected for more than two years [3]. Under the EU AI Act, a system that evaluates learning outcomes is high-risk by classification rather than by anyone's opinion of it [4]. So the design problem was never how to get a model to output a band. It was how to build something an examiner, a regulator and a candidate's appeal can all read. #### The five layers The system listens to the whole interview rather than sampling it, and each layer produces evidence the next one can check. The AI never does the arithmetic. - **1 · Capture**: Transcription with word-level timestamps, hesitations preserved rather than tidied away. A fluency judgement that cannot see where somebody paused is a guess. - **2 · Acoustic measurement**: Speech rate, pause length and position, where hesitation falls, pitch range: measured deterministically in signal-processing code. These are the feature families ETS SpeechRater has used for over a decade [5], and they are arithmetic rather than opinion, so the same audio gives the same numbers every time. - **3 · Pronunciation**: A dedicated phoneme-level model on the Goodness-of-Pronunciation method, trained against a corpus scored by five expert raters. General-purpose multimodal models were tested for this job and were not good enough at phoneme and stress judgement, at 0.21 against 0.61 to 0.74 for specialist models [6]. - **4 · The examiner model**: One pass over the whole interview with every measured value in front of it, applying the official band descriptors. It runs three times, the median is taken, and disagreement between the runs routes the case to a human without anyone asking. It sits behind a swappable interface, on OpenAI audio models with a Gemini fallback, so the instrument does not depend on one vendor staying still. - **5 · Scoring and guardrails**: The four criteria are averaged, rounded by a documented rule and capped where an answer is off topic, all of it in plain code. Integrity checks for memorised answers, read-aloud delivery and synthetic voice raise a review flag; none of them silently alters a score. #### Confidence routing, and the number the design is built around Cambridge publishes the figures that make the case for this shape. Its Linguaskill automarker awarded the same CEFR grade as the examiners on 56.8% of tests marking alone, and on 95.6% under the hybrid model, where a response the computer is not confident about goes to a human [7]. Escalation is what moves that number, not a better model. So the instrument is built around knowing when to stop. Roughly 15% of interviews route to a human reviewer, and a further 5% are audited at random whether the system was confident or not, because a gate tested only on the cases it flagged tells you nothing about the ones it let through. Every reviewed case comes back as labelled data for the next version. #### Fairness and accuracy are release gates These are gates, not results. Stating them the other way round is the failure this whole design exists to avoid. - **Accent fairness**: The pronunciation layer's output is evidence, never the score. Score differences by candidate first language are tested as a formal gate, within 0.10 standard deviations, and a build that fails it does not ship. - **Agreement with examiners**: Quadratic-weighted agreement of at least 0.70 with human examiners, the threshold Williamson, Xi and Breyer set out for automated scoring [8]. - **The validation is scheduled**: The corpus is 300 interviews, each marked by two certificated examiners with a third resolving disagreements. It is booked rather than finished, so this page publishes the gates and no number against them. An accuracy figure without its validation behind it is the first thing a regulator would take apart. #### Where the build stands The pipeline is complete and has been independently verified. Around 1,170 automated tests cover the speaking pipeline, among them a regression suite where disabling any one guardrail breaks specific frozen cases, so a guardrail cannot be quietly dropped without a test going red. The core mathematics was hand-verified against worked examples rather than only against itself. Seven regulator-facing governance documents are drafted, and the rollout is staged behind controls rather than switched on. #### How it started The brief was a frontend. Building it meant living inside the product, and the mock test made the ceiling clear: it could tell a learner whether an answer was right, but not why, and not in a language the learner was comfortable being taught in. We built an AI tutor that teaches English in the learner's own language, and the engagement widened from there into the question the CEO actually had, which was where AI belonged across the company and in what order. #### The estate underneath The engagement is now end-to-end technical ownership of the live production estate: the exam platform, the APIs and the data. On 21 August 2026 the production database held 3,141 live student records, 1,086 of them added in the previous ninety days, and 393 mock assessments completed since November 2025. Those are counts from the system, not a projection. The database was replatformed into Google Cloud's Doha region to meet Gulf data-residency requirements, and the automated test estate across the platform went from around 300 tests to more than 1,400. #### Questions ##### How accurate is the automated IELTS speaking score? No accuracy figure is published for this system. Accent fairness within 0.10 standard deviations by candidate first language, and quadratic-weighted agreement of at least 0.70 with human examiners, are release gates the build has to pass. The validation corpus behind them is 300 interviews, each marked by two certificated examiners with a third resolving disagreements, and it is booked rather than finished. ##### How much of the interview is scored? All of it. The whole interview, eleven to fourteen minutes, passes through five layers of evidence rather than a sample. The AI never does the arithmetic: the four criteria are averaged and rounded by a documented rule in plain code. ##### What happens when the system is not confident? Roughly 15% of interviews route to a human examiner, and a further 5% are audited at random whether the system was confident or not, because a gate tested only on the cases it flagged tells you nothing about the ones it let through. Every reviewed case comes back as labelled data for the next version. #### Sources [1] IELTS test statistics: inter-rater reliability for Speaking: https://ielts.org/researchers/our-research/test-statistics [2] ETS: TOEFL iBT 2026 update, test blueprint and specifications: https://www.ets.org/content/dam/ets-org/pdfs/toefl/toefl-ibt-test-specifications-2026.pdf [3] Tes: Ofqual fines Cambridge English £875,000 over automated marking errors: https://www.tes.com/magazine/analysis/general/ofqual-fines-cambridge-english-ps875000-over-automated-marking-errors [4] EU AI Act, Annex III: high-risk systems, education and training: https://artificialintelligenceact.eu/annex/3/ [5] ETS SpeechRater: automated scoring of spoken responses in the TOEFL iBT test: https://www.ets.org/speechrater.html [6] Exploring the potential of large multimodal models as effective alternatives for pronunciation assessment, arXiv:2503.11229: https://arxiv.org/abs/2503.11229 [7] Cambridge English: Linguaskill, building a validity argument for the Speaking test, June 2020: https://www.cambridgeenglish.org/fr/Images/589637-linguaskill-building-a-validity-argument-for-the-speaking-test.pdf [8] Williamson, Xi and Breyer, A framework for evaluation and use of automated scoring, 2012: https://doi.org/10.1111/j.1745-3992.2011.00223.x Stack: Speech assessment, Signal processing, OpenAI audio models, Gemini, Postgres 16, Google Cloud, Multilingual LLM tutoring, AI roadmap. Last reviewed 22 August 2026. ### AI, data and automation enablement across a global agency (https://www.1aym.com/work/ai-platform-enablement) Reference: D-01. Client: A PE-backed marketing agency. Delivered by 1AYM. Sessions on this agency's AI platform now cost 60% less on average than they did before we re-architected its skills estate. An earlier change, a per-user API call replaced with one ten-minute sync, carries a forecast saving of £2–4M a year, modelled on token consumption and reviewed by the client's finance team. We hold the lead architect role for the platform and work alongside the agency's own finance, systems, data warehouse and AI tooling teams, on an AI programme that is theirs rather than ours. - 60% · Lower cost per session after the re-architecture - £2–4M · Forecast annual saving, finance-reviewed - ~600 · Skills on the platform we contribute to - 80%+ · Skills that exceeded description limits before the rebuild - 4 · Finance systems integrated - Ongoing · Engagement status #### Built with their teams, not around them Worth stating plainly, because case studies routinely blur it: the AI programme at this agency is theirs. Their leadership set the direction, their teams author the skills, and adoption across the business is their achievement. 1AYM is engaged on the engineering underneath, working alongside their finance, systems, data warehouse and AI tooling teams. That is the point rather than a caveat. A supplier who builds in isolation leaves behind a system only they understand, and the client discovers the dependency the first time something breaks after the invoice is settled. Building alongside their people means upskilling them as the work goes. The finance director who shaped the semantic layer with us can defend and extend its definitions himself, and the team maintaining the provisioning system understands why it resolves identity the way it does rather than treating it as a box that must not be touched. The measure of this kind of engagement is not what runs while you are there. It is what the client can still maintain, change and build on once you have gone. #### How the scope grew The engagement started as integration work. It expanded because the problems worth solving kept turning out to sit between teams rather than inside one: a finance reconciliation issue that was really a data contract issue, an identity request that was really a provisioning architecture issue. The value of being forward-deployed is being able to follow a problem across those boundaries instead of handing it off at each one. Over time the role became a bridge between finance, systems, the data warehouse, and the internal AI tooling teams. #### What the work covered - **Finance systems enablement**: Data warehouse and finance systems work across Xero, NetSuite, Zoho and QuickBooks: financial detail, mapping, reconciliation, management-accounts and trial-balance checks, and year-end rollover workflows. - **Identity and provisioning**: Integration pipelines with deterministic identity resolution, dry-run modes, diff-based writes, auditability and safe rollback patterns. - **Internal AI skill platform**: Engineering contributions to an internal Claude Code skill platform that now carries around 600 production skills. The skills themselves are authored across the business by the people who do the work; our part is the plumbing and reliability beneath them. - **Verifier-gated automation**: Managed-agent workflows for finance-critical and high-stakes operational tasks, built so deterministic checks decide what proceeds. - **Grounded analytics**: AI-enabled analytics workflows, including MCP-style patterns connecting the data warehouse and BI layer so AI answers resolve against canonical business metrics. - **Platform reliability**: Production support across Cloud Run and BigQuery/Airflow, deployment fixes, and the ongoing reliability work that keeps daily runs predictable. #### Optimisation nobody asked for The skills library lives in Notion. As written, every user's client fetched from the Notion API directly, which works perfectly at ten users and becomes a problem at a thousand, because the request volume scales with people multiplied by how often they work, against an API with rate limits and an availability budget that was never sized for it. We replaced that with a single synchronisation running every ten minutes, so the platform reads from a local copy rather than every client hitting the source. Request volume stopped scaling with headcount, and the dependency on Notion being fast and reachable at the exact moment someone works went away with it. The forecast saving from that one change is £2–4 million over a year, alongside a system that is materially more reliable. A number that size deserves its provenance: we modelled it on token consumption, and it was reviewed and accepted by a member of the client's finance team rather than asserted by the person who made the change. The arithmetic is not complicated. Under the old pattern, consumption scaled with users multiplied by working sessions multiplied by fetches per session, and at roughly a thousand people that product grows fast, because every one of those fetches pulled content into a context window that someone was paying for. Under the new pattern the cost is fixed: one sync every ten minutes, whatever the headcount is doing. This was not in a brief. It is the kind of thing you only see from inside the system, and the kind of thing worth raising rather than waiting for it to become an incident. #### Phase two: re-architecting the skills estate Sessions on the platform now come back faster and, by our own measurement, cost 60% less on average than they did before the rebuild. Across the number of people using it and the sessions each of them runs in a day, that compounds into a substantial annual saving. 1AYM now holds the lead architect role for the platform, in addition to the sync described above. Phase two happened because phase one worked. Adoption spread, the estate of skills grew with it, and the platform slowed under its own success. The skills had been written as flat, monolithic files, and over 80% of them exceeded the description limits, so every session carried far more context than the task in front of it ever used. The rebuild followed the structure the Anthropic SDK already defines, rather than a shape of our own. - **Skills split to the SDK's structure**: Each skill is split the way the SDK expects, with reference files the model loads on demand instead of one file loaded in full at the start of every session. - **Context injected, not carried**: Custom MCP servers that inject the targeted context a task actually needs, in place of whole files travelling through the session whether or not anything reads them. - **End-to-end sync**: One pipeline from the authoring tool to the running platform, so what an author publishes is what a session gets. - **Simpler to use than it was**: We worked directly with the client's team on how people interact with the platform, so the rebuild left it easier to work with rather than only cheaper to run. #### Why finance work sets the standard Finance-critical automation is the most demanding environment to build in, because there is no acceptable error rate that gets waved through. A reconciliation that is nearly right is a reconciliation that is wrong, and it will be found by an auditor rather than a user. Everything built here inherits that constraint: idempotent operations, dry-run modes, diff-based writes, audit trails, and rollback paths as defaults rather than hardening added later. Stack: Xero, NetSuite, Zoho, QuickBooks, BigQuery, Airflow, Cloud Run, Notion, SCIM, Claude Code. Last reviewed 22 August 2026. ### Identity provisioning for 1,000+ users across 600+ groups (https://www.1aym.com/work/identity-provisioning-at-scale) Reference: D-04. Client: A PE-backed marketing agency. Delivered by 1AYM. When somebody leaves, their access to Notion and Claude ends with their HR record rather than when a colleague remembers. We built the provisioning system that does it from scratch, because the client's IT team was over capacity and off-the-shelf group tooling did not fit a model where one person belongs to many groups: it covers more than 1,000 users across over 600 groups and has run for seven months. A membership matrix that size was never going to be kept right by hand. - 1,000+ · Users provisioned - 600+ · Groups managed - 7 months · Running in production - HR system · Single source of truth #### The problem Two tools needed access managed across the whole organisation, against a membership model where one person belongs to many groups and group membership changes as people join, move and leave. Done manually, a matrix of that size is not merely tedious. It is unreliable in a specific and dangerous way. Manual provisioning fails safe on joiners, because someone complains when access is missing. It fails unsafe on leavers, because nobody complains about access that should have been removed. That asymmetry is how organisations accumulate active accounts belonging to people who left months ago. #### Why it was built rather than configured The internal IT team was over capacity and did not have the time to take it on, which is the ordinary reason this kind of work does not get done anywhere. Off-the-shelf group tooling was considered and rejected. The organisation's membership model (overlapping groups, HR as the source of truth, two target systems with different provisioning semantics) did not fit the shape those tools assume, and bending the model to fit the tool would have left the mismatch to be handled manually anyway. #### How it works - **HR as the source of truth**: Provisioning is driven from the HR system, so joining, moving and leaving are reflected without a separate administrative step that someone has to remember. - **Deterministic identity resolution**: Matching a person across systems is done by explicit rules rather than fuzzy heuristics, because a wrong match grants the wrong person access. - **Diff-based writes**: Each run computes the difference between intended and actual state and applies only that, so a run is idempotent and a retry cannot compound. - **Dry-run mode**: Changes can be inspected before they land, which is necessary when a single run can alter access for hundreds of people. - **Audit trail and rollback**: What changed, for whom, and why, with a defined path back from a bad run. #### Outcome The system has run for seven months. Access reflects the HR system rather than someone's memory, joiners are provisioned without a ticket, and leavers lose access because a record changed rather than because somebody noticed. This was built inside the wider enablement engagement at the same organisation, and shares its engineering defaults: idempotency, dry runs, auditability and rollback are the baseline rather than additions. Stack: SCIM, Notion, Claude, HR system integration, Deterministic identity resolution. Last reviewed 27 July 2026. ### A regulated product that was not going to ship, shipped. (https://www.1aym.com/work/clinical-trials-platform-rescue) Reference: D-05. Client: A clinical-research software start-up. Delivered by 1AYM. A clinical-research software start-up had a part-built platform, no engineering environment its own team could ship from, and a launch that was not going to happen on the path it was on. We worked with the founders: first a development environment, with repositories, environments, continuous integration, review gates and a release discipline, then completion of the platform, then delivery to launch alongside their team. The product is live and in market, and the environment it ships from belongs to the company rather than to us. - Founders · Engagement partner - Dev environment · Built for the team to ship - Completed · Platform modules brought to release - Live · Launched and in market #### Where it was The company had built part of a product and could not get it out. The platform existed in pieces, there was no engineering environment that could carry a change from somebody's branch to a release anyone would trust, and on the path it was on it was not going to launch. That is a common shape and it is rarely a shortage of code. What is missing is the ordinary machinery a regulated product needs before a release can be a routine event: somewhere for the work to live, environments to run it in, checks that run on every change, and a repeatable way to put a version in front of users. Without it, every release is an act of nerve, and a team that is nervous about releasing stops releasing. #### What we put in place In this order, because none of the later pieces is safe without the first one. - **A development environment**: Repositories, environments, continuous integration and review gates, so a change is written, reviewed, tested and promoted the same way every time. Built for the founders' own team to work in rather than for us to operate on their behalf. - **Release discipline**: How a version is cut, what has to pass before it moves, and what happens when something fails: defined, and in use. Discipline that holds only while the supplier is in the building is not discipline. - **Completion of the platform**: With a path to release in place, the remaining modules were built out and taken to a state that could ship rather than to a state that could be demonstrated. - **Delivery to launch**: We took the platform to launch with the founders rather than handing over a repository and a document. A first release is where everything the process missed turns up, so it is the release worth being in the room for. #### Why it matters in regulated research Clinical research software is not judged only on whether it works. It is judged on the record it leaves behind: who did what, under whose delegated authority, against which version of which document, and whether all of that can be produced on request. The platform holds those systems on one inspection-ready record rather than in separate products that have to be reconciled afterwards. Those systems are trial management (CTMS), the electronic trial master file (eTMF), the investigator site file and the pharmacy site file (eISF and ePSF), quality management (eQMS), corrective and preventive actions (CAPA), and digital delegation of authority. It is built for NHS and hospital R&D, academic research, contract research organisations and mid-market sponsors, and its documentation is aligned to MHRA, FDA and EU Annex 11 expectations. That is the product's own public description of what it is for, and this page repeats it at that level. #### What the founders are left with The product is live and in market. A launch that was not going to happen on the previous path happened on this one, and it happened in a way the company can repeat. What stays behind counts as much as what shipped. The development environment is theirs and their own team runs it, the release discipline is theirs and in use, and the next release goes out through the same path that shipped this one. That is the measure we apply to every engagement on these pages: not what runs while we are there, but what the client can maintain, change and build on once we have gone. Stack: Development environment, Continuous integration, Review gates, Release management, CTMS, eTMF, eQMS and CAPA, Regulated documentation. Last reviewed 22 August 2026. ## Reference ### What is an AI harness? (https://www.1aym.com/what-is-an-ai-harness) Definition: The operating environment a model vendor builds and maintains with the model, on which an organisation places its own architecture and skills packages. Claude Code, Codex, Cursor. That's the harness. The vendor builds it. You still have to own the architecture and skills that sit on it. Those three are the operating environments Anthropic, OpenAI and Cursor ship and maintain with the model. What the company still writes is the architecture and the skills packages that sit on that harness: how work is structured, what the model is allowed to do, and how changes reach production. 1AYM is a consultancy that designs that layer so an organisation can use first-party vendor harnesses without being locked to one vendor's stack. #### Who builds the harness The harness is the vendor's product. Anthropic maintains Claude Code. OpenAI maintains Codex and the ChatGPT tooling around it. Cursor maintains the harness that ships with its editor. Each one is updated on the vendor's schedule, against the vendor's model, with the vendor's evaluation and support behind it. A company that sets out to build a competing harness is taking on that maintenance load as well as the work that sits on it. The useful question is not how to replace Claude, ChatGPT or Cursor. It is what has to sit on those harnesses so the organisation can use more than one of them, and so the second team does not start from a blank page. #### What sits on a harness The vendor harness runs the model and the tools it ships with. It does not know the company's permission model, its definitions, or the way its engineers are allowed to ship. That layer is architecture and skills packages, written once and reused across the first-party harnesses the teams already have. - **Architecture**: How work is split, which harness is used for which job, and where the company's systems of record sit relative to the model. - **Skills packages**: Reusable instructions, tools and context that a vendor harness can load, so a workflow is not rewritten for each team and each tool. - **Guardrails**: What the model may read and write, on whose behalf, and which actions need a person. - **Route to production**: Review, CI and the same path every other change already takes, so AI-assisted work is not a side channel. #### What an AI harness is not It is not a product 1AYM sells, and it is not a reason to leave Claude, ChatGPT or Cursor. Those vendors already ship the harness. Replacing them with a house-built equivalent means owning the runtime as well as the work that sits on it. It is also not a framework you install. The architecture, the skills packages and the permission model encode decisions specific to your business. Frameworks help you wire calls; they are not the environment the model vendor maintains, and they are not a substitute for it. #### When the architecture matters The first team can work directly in Claude Code, Codex or Cursor and ship. By the third, the same integrations, permissions and review rules are being rewritten, and nobody can say whether the three setups agree. The other common trigger is work that touches money, identity or a regulated process, where the vendor harness is the right runtime and the company still has to specify what may proceed. The cost of leaving the skills layer unowned is measurable rather than theoretical: on one platform we hold the lead architect role for, over 80% of skills had grown past the description limits before the rebuild, and re-architecting the estate cut the average cost per session by 60%. That is skills debt, priced. #### Questions ##### Is an AI harness the same as an agent framework? No. An agent framework is a library for constructing agents: orchestration primitives, tool-calling conventions, memory abstractions. An AI harness is the environment the model vendor ships and maintains. A framework might run inside a harness, or beside one; it is not a substitute for Claude Code, Codex or Cursor, and it is not the architecture and skills packages the company still has to write. ##### Do we need to build our own harness? Almost never. The vendors already maintain one. The work is to put architecture, skills packages, guardrails and CI on the harnesses the engineers are already using, so the company is not rewriting that layer for each tool and each team. Building the runtime first, in place of Claude, ChatGPT or Cursor, tends to produce a second maintenance burden the business never needed. ##### Does using a vendor harness lock us into one model provider? It can, if the architecture and the skills are written for one tool only. Treating those as a layer that sits on the harness, rather than inside one vendor's product, means the same packages can be aimed at Claude Code, Codex or Cursor as the work requires. Model choice stays a decision; it does not become the platform. Last reviewed 29 August 2026. ### Agent proposes, verifier gates (https://www.1aym.com/agent-proposes-verifier-gates) Definition: An operating pattern for dependable AI automation in which the AI performs flexible reasoning while deterministic validators, schemas, tests, and human review gates decide whether the output is safe to proceed. “Agent proposes, verifier gates” is an operating pattern for dependable AI automation: the AI performs the flexible reasoning (research, drafting, transformation, generation) while deterministic validators, rules, schemas, tests, and human review gates decide whether the output is safe to proceed. The model is never the last thing to touch a decision. It is what makes automation dependable in finance, data migration, and compliance workflows where “mostly correct” is not good enough. #### Why “mostly correct” fails A model that is right 95% of the time sounds excellent until you run it a thousand times a day against a ledger. Then it is fifty wrong entries a day, arriving with the same confident tone as the correct ones, in a system where the cost of a wrong entry is not one-twentieth of the value of a right one. Accuracy is the wrong frame for automation. The question is not how often the model is right. It is what happens on the occasions it is wrong, and whether the system notices before the consequence lands. #### How the pattern works The work is split by what each component is actually good at. Models are good at ambiguity, language, and synthesis. Deterministic code is good at being certain. The pattern uses each for what it is reliable at, and never asks the model to certify itself. - **The agent proposes**: It reads the ticket, drafts the entry, maps the fields, extracts the values, or writes the summary: the part that requires judgment across messy inputs. - **The verifier gates**: Schema validation, business rules, reconciliation against a source of truth, type and range checks, and tests. None of it involves a model. All of it can be reasoned about and unit tested. - **Pass proceeds, fail holds**: Passing output continues automatically. Failing output routes to human review with the reason attached, rather than being silently retried or silently shipped. - **The gate is the audit trail**: Because every output passes an explicit check, the record of what was checked and why it passed is a by-product of running the system, not extra compliance work bolted on afterwards. #### What the verifier actually checks The strongest verifiers are boring. They assert that a total reconciles, that an identifier exists in the system of record, that a date falls inside the period, that a required field is populated and typed correctly, that a proposed write is idempotent. Where a deterministic check is genuinely impossible, the gate becomes a human one, but a narrow one, presented with the specific claim to confirm rather than the whole task to re-do. The aim is to spend human attention only where it is irreplaceable. #### Where the pattern applies It earns its keep anywhere the cost of a silent error exceeds the cost of a held item: finance-critical automation, data migration validation, knowledge extraction, semantic-layer query validation, report generation, operational triage, and internal tooling that writes to systems of record. It is unnecessary where output is inherently reviewed by the person who asked for it: a drafting assistant a human reads before sending does not need a gate, because the human is one. #### Questions ##### Does this slow the system down? Not meaningfully. A deterministic check costs microseconds against a model call's seconds. What changes is that some items stop instead of completing, which is the job. ##### Can the verifier be another model? Sometimes, but it is a weaker guarantee and should never be the only gate on a consequential action. A model checking a model shares failure modes with it, and both can be confidently wrong about the same thing. Model-based evaluation is useful for scoring quality at scale; deterministic validation is what you use to decide whether something is safe to write. ##### How is this different from a human-in-the-loop workflow? Human-in-the-loop usually means a person reviews everything, which does not scale and quietly degrades into rubber-stamping. This pattern routes only what fails an explicit check to a person, with the reason attached, so human attention is spent on genuine exceptions rather than spread thin across a queue that is mostly correct. Last reviewed 27 July 2026. ### How to choose an enterprise AI consultancy in the UK (https://www.1aym.com/how-to-choose-an-enterprise-ai-consultancy) Definition: A consultancy that takes an organisation from AI strategy to governed systems running in production, combining architecture, hands-on engineering and implemented controls rather than advice alone. Choosing an enterprise AI consultancy in the UK comes down to four inspectable checks: who writes the code, one production workflow with measurements and their provenance, governance implemented as controls rather than documents, and what the client keeps when the engagement ends. #### Four checks, and how to read a pitch 1. **Ask who writes the code**. Meet the people who would deliver, not the people who sell. Firms where the same people write the strategy and the code cannot hide behind a handover. 2. **Ask for one workflow in production**. Ask for one real workflow taken to a system people use daily, with the measurements and their provenance stated. A demonstration is not that proof. 3. **Ask whether governance is implemented as controls**. Identity, permissions, deterministic checks and human gates are inspectable. Policy documents and committees are not. 4. **Ask what you keep when they leave**. The client should own the code, the architecture and the operating model, with a written handover, and be able to continue without the consultancy. | | What a pitch looks like | | --- | --- | | Ask who writes the code | A partner sells the work and a separate bench builds it. | | Ask for one workflow in production | A slide deck, unused seats, or a figure with no source. | | Ask whether governance is implemented as controls | A governance workstream that produces a policy set and a steering committee. | | Ask what you keep when they leave | A change-request route back to the firm, or knowledge that leaves with the team. | #### What production proof looks like, from work we have shipped These four checks are answerable from production systems, not from a demonstration. The figures below are already published on our engagement files, each with its provenance. We are based in Stoke-on-Trent and deliver across the UK, the Gulf and the US. - **Who writes the code**: At 1AYM the people who write the strategy write the code. Every engineer is certified on the platforms we build on or has shipped inside a top-tier engineering organisation. Meta, Spotify, UBS, Starling Bank, S&P Global and Sky are former employers of our engineers, not clients. - **One production workflow**: A government-accredited EdTech's speaking assessment runs as a measuring instrument on the live estate: 3,141 live student records on 21 August 2026, counts from the production database, and an automated test estate grown from around 300 tests to more than 1,400. A PE-backed marketing agency's platform sessions cost 60% less on average after we re-architected its skills estate, by our own measurement. Identity provisioning for 1,000+ users across 600+ groups has been running for seven months. - **Governance as inspectable controls**: The pattern we implement is the model proposing and something deterministic deciding: an agent drafts, and validators, reconciliations and human gates decide what proceeds. Evaluation runs in CI. We hold no ISO or SOC certification; controls are evidence, not a substitute for a certificate. - **What the client keeps**: Architecture, code and operating model are documented and handed over. Nothing depends on us staying. - **Vendor status is disclosure, not proof of fit**: 1AYM is an OpenAI Select Partner: a company status, not a certification, and not evidence of fit by itself. The founder holds Claude Certified Architect – Professional and Claude Certified Associate – Foundations, independently verifiable on Credly. 1AYM holds no partner status with Anthropic. Another client environment reached 91% adoption in North America, a figure published in the Anthropic customer case study. #### The market is four different businesses wearing one label Search for an AI consultancy in the UK and the results mix four kinds of firm that share a label and almost nothing else. Which one you need depends on the problem, and most bad engagements start by buying the wrong kind, not the wrong firm. - **The Big Four and large systems integrators**: Built for board-sponsored, multi-year transformation with global delivery and regulatory cover. Genuinely good at programme management at scale. The trade-off is a partner who sells, a bench that builds, and a cost structure that needs a large programme to justify itself. - **Specialist consultancies**: Senior practices where the people who scope the work deliver it. Suited to organisations that want a production system, a working operating model and their own capability at the end, without funding a programme office. - **Staff augmentation**: Sells individual engineers by the week. Useful when you already have the architecture and the leadership and simply need hands. It transfers no design responsibility: if the system is wrong, that is your problem. - **Automation agencies**: Assemble chatbots and workflow tools quickly for SME budgets. The right answer for small internal conveniences, and the wrong one for anything that touches a ledger, a regulator or a system of record. #### What the market actually charges, from the only prices firms publish by name No UK body publishes a genuine benchmark of consultancy fees. The Management Consultancies Association publishes market size rather than prices, Source Global Research publishes pricing sentiment with the detail behind a paywall, and the Crown Commercial Service publishes its consultancy framework's grade structure and the fact of contractual caps without the figures. The fee guides that circulate online cite nobody, which is why their numbers disagree with each other. The only public, firm-attributable prices in the UK are the per-day fee schedules suppliers must publish to sell through the government's G-Cloud framework. G-Cloud 14 is in force until October 2026, with a successor announced in August 2026, and most of its cards were priced in April and May 2024. Read them as a floor rather than a quote: ONS figures put producer price inflation for professional, scientific and technical services at 3.5% in the year to the first quarter of 2026. Suppliers price against SFIA, a seven-level seniority scale, which is what makes the cards comparable. The figures exclude VAT and assume an onshore person-day with professional indemnity cover included. At level 5, the senior specialist grade, published prices run from £800 to £2,380 a day across the eleven supplier fee cards read for this page on 1 September 2026, with a median of £1,500. The large firms sit at the top of that level: Accenture, Deloitte and KPMG publish £1,580 to £2,380 at level 5 and £1,880 to £2,855 at principal or partner level. Practices outside the large firms publish £800 to £1,400 for the same level, with Fusion AIX at the floor of that band. Whole engagements carry fixed prices too. On its G-Cloud 14 card Gartner prices whole engagements rather than days, publishing £70,200 for AI use case elicitation and prioritisation, £142,400 for an AI capability maturity assessment and £203,500 for an AI strategy and roadmap. What surprises buyers in the catalogue is that the large firms do not simply price higher. Across whole listings their published bands are wider at both ends, typically holding lower floors as well as higher ceilings, and the floor is where the misreading happens. None of which makes a published figure a quote. The useful question is not the price of a day but what a fixed scope delivers and who carries the risk when the work runs long. 1AYM sells three shapes: fixed-scope statements of work, a retained architecture engagement, and engineers embedded in a client's own programme. None of them is priced on this page, and the embedded engagement states its commercial terms on its own page. Four things in the same public record are worth more to a buyer than the bands themselves. - **The headline low is not the specialist**: Accenture's G-Cloud listing spans £95 to £2,240 a day. Its own fee card explains the spread: £95 buys a level 1 junior working offshore, while the onshore senior specialist at level 5 is £1,580. A quoted figure carries no information until you know who does the work and from where. - **What a person is paid is not what a firm charges**: Advertised interim figures for individual AI and machine learning specialists sit at a median of £560 to £600 a day, depending on the specialism (IT Jobs Watch, six months to September 2026). That is a multiple below firm pricing at comparable seniority, and the gap is overhead, cover, insurance and margin. A firm quoting at individual level is worth interrogating on what carries the delivery risk. - **Paying for time is not paying for an outcome**: The National Audit Office, reviewing government spending on external consultants, found that input-based pricing “often fails to deliver the best value for money”, and that outcome-based payment may offer better value where the outcomes are clearly measurable. That is third-party support for buying a fixed scope over open-ended time. - **Nobody publishes the independent tier**: There is no public, primary source for what a senior specialist working alone charges a client: G-Cloud publishes no headcount, and recruitment guides measure what a person is paid rather than what a client pays. Bands for that tier circulate widely and cite nothing, so this page does not invent one. #### Ask who writes the code The single most predictive question in selection. Firms where strategy and engineering are separate departments produce strategies that engineering quietly rewrites, and systems that drift from the deck that sold them. Firms where the same people write the strategy and the code cannot hide behind the handover, because there is none. Ask to meet the people who would deliver, not the people who sell. Ask what they personally shipped in the last quarter. A practice that cannot answer that question with named systems is a broker. #### Ask for one workflow in production, then read the proof properly Demonstrations are cheap and production is expensive, so production is the evidence that matters. Ask every candidate for one real workflow they took from ambition to a system people use daily, and how long it took. Then read the numbers with their provenance attached. A figure published by the client, a figure the consultancy measured, and a forecast are three different strengths of evidence, and a firm that blends them into one claim is telling you how it will report your programme too. #### Governance you can inspect beats governance you can read Every firm will say the word governance. The separating question is where it lives. If the answer is a policy document and a committee, the controls depend on people remembering them. If the answer is identity and permissions the system enforces, deterministic checks that gate what an agent may do, evaluation that runs on every change and human approval on consequential actions, the controls hold when nobody is watching. The pattern to look for is the model proposing and something deterministic deciding: an agent drafts, and validators, reconciliations and human gates decide what proceeds. A consultancy that cannot describe its gating pattern in that level of detail has not built one. #### Vendor position: partner status is disclosure, not proof of fit Most AI consultancies hold some vendor relationship, and the honest ones state it plainly, because it shapes advice. A reseller attached to one vendor will recommend that vendor. A model-selective practice chooses by workload and production evidence, and can show you systems running on more than one stack. Partner statuses are worth reading precisely. 1AYM, for example, is an OpenAI Select Partner: a company status that describes a relationship with the vendor, stated so a buyer can weigh it, not evidence by itself that the firm fits your problem. Treat any partner badge the same way, from any firm: as a disclosure to interrogate, not a shortcut past the four checks above. #### Ask what you keep when they leave The end state of a good engagement is that the client owns the code, the architecture, the operating model and the capability, and could continue without the consultancy. Ask directly: what do we own on the last day, who can run it, and what does it cost to keep running. Firms whose commercial model depends on you being unable to leave will resist that question. That resistance is the answer. #### When a consultancy is the wrong answer If the problem is a well-served product category, buy the product. If you have strong engineering leadership and a clear architecture, hire or borrow engineers instead of buying design you already have. If nobody senior owns the outcome internally, fix that first, because no external firm can substitute for an absent owner. A consultancy earns its fee where strategy, architecture and engineering have to move together and the organisation wants to keep the result: taking the first governed workflow into production, standing up the platform and controls around it, and transferring the capability to run it. #### Questions ##### What four checks should you use to choose an enterprise AI consultancy in the UK? Choosing an enterprise AI consultancy in the UK comes down to four inspectable checks: who writes the code, one production workflow with measurements and their provenance, governance implemented as controls rather than documents, and what the client keeps when the engagement ends. ##### What does production proof look like for an enterprise AI consultancy? Production proof is one real workflow people use daily, with measurements and their provenance stated. A client-published figure, a consultancy measurement and a forecast are three different strengths of evidence. A demonstration is not production. ##### Should we choose a Big Four firm or a specialist consultancy? Match the firm to the shape of the work. A board-sponsored, multi-country transformation with heavy regulatory reporting suits a large integrator. A first production system, a platform build or an engineering-led adoption programme suits a senior specialist practice, because you are buying judgement and code rather than programme management. ##### What does a partner status with OpenAI or Anthropic actually tell you? It tells you the firm has a formal relationship with that vendor, which usually means earlier access, direct support channels and co-delivery experience. It does not tell you the firm is right for your problem, and it is worth asking any partner firm to show work delivered on stacks outside that vendor. ##### How quickly should work reach production? For a first governed workflow, weeks to a small number of months, not quarters. The honest constraint is usually access, data and approvals rather than engineering. A firm that cannot name what it would ship in the first month is planning a long discovery. ##### What should the contract say about ownership? That the client owns the code, the architecture, the documentation and the operating model, with no licence back to the consultancy required to keep running the system. Handover, runbooks and capability transfer should be deliverables with acceptance criteria, not goodwill. ##### Does the consultancy need to be UK-based? For UK organisations with data-residency, procurement or sector-regulatory constraints, a UK-headquartered practice working to UK GDPR and the Data Protection Act 2018 simplifies the compliance conversation. 1AYM is based in Stoke-on-Trent and delivers across the UK, the Gulf and the US. What matters more than the address is whether the firm architects for your residency requirements and can evidence it in production. ##### How much does an enterprise AI consultancy cost in the UK? No UK body publishes a fee benchmark, so the only public, firm-attributable prices are the per-day fee schedules suppliers must publish to sell through the government's G-Cloud framework. On the G-Cloud 14 cards, a senior specialist at SFIA level 5 runs from £800 to £2,380 a day with a median of £1,500: £1,580 to £2,380 at the large firms, £800 to £1,400 at practices outside them. Whole engagements carry fixed prices too, with Gartner's card publishing £70,200 for AI use case prioritisation up to £203,500 for an AI strategy and roadmap. A headline low is not comparable until you read the fee card behind it: Accenture's listing starts at £95 a day, which its own card shows is a level 1 junior offshore, against £1,580 for the onshore senior specialist. Most cards were priced in April and May 2024 and are a floor rather than a quote. These are other firms' published prices, not 1AYM's; 1AYM's engagements are not priced on this page, and the question that decides value is what a fixed scope delivers and who carries the risk. #### Sources [1] Crown Commercial Service, G-Cloud 14 framework RM1557.14 (2024): https://www.gca.gov.uk/agreements/RM1557.14 [2] Accenture (UK), G-Cloud 14 data science services listing (2024): https://www.applytosupply.digitalmarketplace.service.gov.uk/g-cloud-14/services/427074791998071 [3] Accenture (UK), G-Cloud 14 SFIA fee card (2024): https://assets.applytosupply.digitalmarketplace.service.gov.uk/g-cloud-14/documents/92191/570124328858333-sfia-rate-card-2024-04-21-0313.pdf [4] Deloitte, G-Cloud 14 specialist pricing document (2024): https://assets.applytosupply.digitalmarketplace.service.gov.uk/g-cloud-14/documents/92485/705246728574904-pricing-document-2024-04-25-1459.pdf [5] KPMG, G-Cloud 14 SFIA fee card (2024): https://assets.applytosupply.digitalmarketplace.service.gov.uk/g-cloud-14/documents/93303/654229059113914-sfia-rate-card-2024-04-22-1527.pdf [6] Faculty Science, G-Cloud 14 pricing document (2024): https://assets.applytosupply.digitalmarketplace.service.gov.uk/g-cloud-14/documents/703150/248323395630555-pricing-document-2024-05-03-1630.pdf [7] Fusion AIX, G-Cloud 14 SFIA fee card (2024): https://assets.applytosupply.digitalmarketplace.service.gov.uk/g-cloud-14/documents/721839/404114134907556-sfia-rate-card-2024-05-07-0614.pdf [8] Gartner, G-Cloud 14 pricing document (2024): https://assets.applytosupply.digitalmarketplace.service.gov.uk/g-cloud-14/documents/92542/280137345185975-pricing-document-2024-05-02-1422.pdf [9] National Audit Office, Lessons learned: the government's use of external consultants, HC 1381 (2025): https://www.nao.org.uk/wp-content/uploads/2025/11/governments-use-of-external-consultants.pdf [10] Office for National Statistics, Producer price inflation, UK: March 2026, including services January to March 2026 (2026): https://www.ons.gov.uk/economy/inflationandpriceindices/bulletins/producerpriceinflation/march2026includingservicesjanuarytomarch2026 [11] IT Jobs Watch, Artificial intelligence: UK interim market pay (2026): https://www.itjobswatch.co.uk/contracts/uk/artificial%20intelligence.do Last reviewed 1 September 2026. ### How to roll out Codex and Claude Code across an engineering organisation (https://www.1aym.com/codex-and-claude-code-enterprise-rollout) Definition: The organisational work of putting shared skills packages, a permission and identity model, review gates and CI evaluation around vendor coding agents such as Codex and Claude Code, so agent-written changes reach production by the same path as every other change. Rolling out Codex and Claude Code across an engineering organisation is four pieces of work, and buying licences is none of them. Skills packages come first: the repository conventions, tools and context an agent loads before it does anything, owned once and pointed at whichever vendor harness a team uses. A permission and identity model decides which repositories, systems and credentials an agent may reach, on whose behalf, and which actions stop for a person. Agent-written changes then go through the review path every other change already takes, with their provenance recorded on the change. And the skills themselves are evaluated in CI, because a skill that has drifted from the codebase is worse than no skill at all. Run it team by team, starting with one repository that already ships often, and read adoption from merged changes and rework rate rather than seats activated. #### What a rollout actually consists of A licence gives every engineer a coding agent and nothing else. The layer around it is what decides whether the third team gets the same agent the first team did: the engineering below, and one person answerable for it. - **Skills packages**: How this codebase is tested, what a good change looks like here, which internal services may be called and how. Written once and loaded by Codex, Claude Code or Cursor as the work requires, rather than reinvented per team. - **Permissions and identity**: Which repositories, cloud accounts and credentials an agent reaches, and on whose behalf. Agent identity is the part most rollouts postpone and the part that decides what a mistake can touch. - **Review gates**: The same pull request, the same code owners and the same required checks as every other change, with provenance recorded on the change so nobody has to guess later which ones an agent wrote. - **CI evaluation**: The skills are tested like code. A package that has drifted from the repository fails a check, rather than quietly producing worse changes for a quarter. - **A named owner**: One person accountable for the skills estate and the permission model. Without that, everything above becomes nobody's job by the third team. #### The order matters more than the tool choice Start in one repository that already ships frequently and has tests worth trusting, because it is the only place you can tell whether the agent helped. Get skills and permissions right there, in that order: skills give the agent enough context to be useful, permissions decide what its mistakes can reach. Review gates come next, before the second and third team rather than after them. CI evaluation of the skills is the piece most organisations defer, and it is the one that decides whether the estate still works in six months. The trade-off is honest: done in this order the first three weeks look slow, because the visible output is configuration rather than merged pull requests. A rollout that starts with breadth instead buys three weeks and then spends a quarter unpicking a hundred private conventions formed in the first fortnight. #### Where do-it-yourself rollouts stall None of this is beyond a competent in-house team, and plenty of organisations do it themselves. The ones that stall tend to stall in the same places, and every one of them is organisational rather than technical. - **Skills nobody owns**: Each team writes its own, none are reviewed, and nobody can say whether the setups agree. The estate grows faster than the ability to reason about it. - **Permissions granted per person**: Access follows whoever asked, so what an agent can reach is whatever its operator happens to hold, and no one can answer what a bad run would have touched. - **A side channel to production**: Agent-written changes attract a lighter review because they read tidily. Tidy is the failure mode: fluent code passes eyes that would have caught the same error written badly. - **Adoption measured by licences**: Seat counts rise, nobody can name a change that reached production because of the tools, and the renewal conversation has no evidence in it. #### An unowned skills estate accrues debt you can price Leaving this layer unowned has a price, and on one engagement we can put a measurement against part of it. On an enterprise platform we hold the lead architect role for, the skills estate had drifted far enough that re-architecting it produced a material reduction in the average cost per session. The engagement file records the measurements and their provenance; both are 1AYM measurements on that platform, taken before and after the change, not a benchmark and not a forecast. The mechanism was dull, which is the point. Skills written as flat, monolithic files dragged far more context into every session than the task in front of them used. Nobody had decided that; it accumulated, one unreviewed package at a time, which is exactly what an unowned estate does. The fix was architecture rather than heroics: routing, description budgets and evaluation, the same layer this page describes. #### How do you measure adoption honestly? Seat activation measures procurement. It tells you what was bought, not what changed. A handful of numbers, read per team and against that team's own baseline, tell you whether the rollout worked. - **Share of merged changes with agent involvement**: Recorded on the change automatically, never self-reported. A survey of how useful engineers found the tools measures enthusiasm. - **Time from first draft to review-ready**: The interval these tools genuinely compress. If it has not moved, the skills packages are too thin to carry the codebase. - **Reviewer time per change**: If this is rising, the rollout has moved cost from author to reviewer rather than removing it. That is a real result and it belongs in the report. - **Rework within thirty days**: Changes reverted or substantially rewritten. The number that says whether faster output was worth having. #### When not to roll out Do not roll out to an organisation that cannot merge a small change quickly today. Agents raise the volume of proposed changes, so if review is already the bottleneck, more proposals make it worse and the tools take the blame for a queue that predates them. Do not roll out where nobody senior owns the outcome. Skills, permissions and gates are decisions with trade-offs attached, and a rollout without an owner reverts to whatever each team improvises. And do not buy a programme where one squad and a fortnight would settle the question. Put Codex or Claude Code on a real repository with two weeks and a target, and count the changes that merged. If the answer there is no, an organisation-wide rollout will not rescue it, and the money is better spent on the tests and the review capacity that would have made the answer yes. #### Questions ##### Should we standardise on one coding agent, or run Codex and Claude Code side by side? Run both, and write the layer around them once. The agents differ in harness behaviour and in what they are strongest at, and teams will form preferences you cannot usefully overrule. What must not differ is the skills packages, the permission model and the review path, because those are what make the estate legible. Standardise the operating layer, keep the tool a team-level choice, and revisit it when a measured difference in the numbers appears. ##### Do we pilot with one team first, or roll out to all of engineering? One team, one repository that ships frequently, and a fixed end date. A pilot exists to produce evidence a wider rollout can be designed from: which skills the codebase actually needs, where permissions bite, what review load looks like. Rolling out to everyone first inverts that, and the conventions formed in the opening fortnight are then the thing you have to unpick before anything can be shared. ##### How long before engineers feel the difference? Individuals feel it in days, because the tools are useful on day one without any of this. The organisation feels it when skills packages carry real repository context and the review path has stopped being a queue, which is usually six to twelve weeks depending on how much of the permission and identity model already exists. Anyone promising an organisation-wide change inside a fortnight is describing licence activation. ##### Who owns the skills packages once the rollout is done? The client, in the client's repositories, reviewed like any other code, with a named owner for the estate and a maintainer per package. If an external firm holds the skills, the organisation has bought a dependency instead of a capability. The test is simple: can your engineers change a skill, evaluate it in CI and ship it without anyone external in the loop. ##### What do we do about engineers already using these tools without a policy? Treat that as the starting material, not as a violation. Find out which repositories, which tools and which credentials are already in play, then write the permission model and the skills packages around what is genuinely happening. A ban produces the same usage with less visibility, and the conventions those engineers have already worked out are usually the best first draft of the skills estate. Last reviewed 31 August 2026. ### Gulf AI data residency: the architecture questions buyers should ask (https://www.1aym.com/gulf-ai-data-residency) Definition: The set of architecture decisions that determine where data governed by a given jurisdiction is stored, whether the in-country cloud region runs the model that processes it, and where inference on it actually takes place. AI data residency in the Gulf is three separate questions that buyers usually collapse into one. First, where the law actually requires data to stay. As of September 2026, and as a reading of the instruments rather than legal advice: in Saudi Arabia, the UAE and Qatar the personal-data statutes regulate cross-border transfer rather than imposing general localisation, and the hard in-country rules sit in sectoral instruments such as the Saudi government cloud policy, UAE health and banking rules, and the Qatar Central Bank's cloud regulation. Second, whether the cloud region runs the model at all: a live region in-country can host the AI platform while serving few or none of the models in-region. Third, where inference actually happens, because data at rest is not data in use, and the managed routes to frontier models from Gulf regions are mostly globally routed by design. A buyer who asks only whether there is a region in-country gets a compliant-looking architecture that quietly processes prompts on another continent. #### Three questions, and only one of them is legal Three different things get called data residency in the same sentence, and they have different answers. The first is legal: what an instrument actually obliges you to keep inside the country. The second is infrastructural: whether the cloud region you chose runs the model you want. The third is operational: where the computation on your prompt happens, which is not settled by where your database sits. Buyers usually ask a version of the first question, accept a version of the second as the answer, and never ask the third. This page covers the GCC jurisdictions where that pattern causes the most damage: Saudi Arabia, the UAE and Qatar, with Bahrain relevant as a region rather than as a separate legal analysis. - **The legal question**: Which instrument binds you: the personal-data statute, a sector regulator's rulebook, a government cloud policy, or a free-zone regime. They give different answers, and the sectoral ones are the strict ones. - **The region question**: Whether a live region exists in the country, and separately whether that region serves the model you intend to call. As of September 2026 those are not the same fact in any Gulf country. - **The inference question**: Where the prompt is processed and the output produced, as distinct from where the data is stored. Vendor documentation answers this explicitly, and the answer is often not the region you selected. #### Where does Gulf law actually require data to stay? Start with what the personal-data laws do not say. Saudi Arabia's PDPL, the UAE's Federal Decree-Law 45 of 2021 and Qatar's Law 13 of 2016 all regulate cross-border transfer: conditions, safeguards, assessments. None of them contains a general requirement that personal data be kept inside the country. If your residency requirement came from a summary of one of those laws, it is probably stricter than the law. The binding localisation rules are sectoral, and where they apply they are strict. What follows is a reading of primary instruments as of September 2026, not legal advice, and each of them should be re-verified against the sources listed below and with your own counsel before it becomes a design decision. - **Saudi government data**: The MCIT Cloud First Policy states that all data in both the Government Cloud and the Commercial Governmental Cloud should be located geographically inside the borders of Saudi Arabia. The PDPL, enforced by SDAIA since September 2024, instead governs transfer: standard contractual clauses, binding common rules or accredited certification, with a transfer risk assessment in defined cases, no general prior approval requirement, and no adequacy list published as at September 2026. - **Qatar, which runs the other way**: The MCIT Cloud First Policy P005 classifies government data from C0 to C4, works through an endorsed provider list, and keeps the C4 tier off cloud entirely. The PDPPL pushes in the opposite direction: Article 15 forbids a controller from restricting cross-border flow except where the processing would breach the law or risk serious harm, and it sits in the penalty schedule. Qatar is the Gulf jurisdiction whose data-protection statute argues against localisation rather than for it. One sectoral exception is the Qatar Central Bank, whose cloud computing regulation, in force since April 2024, requires licensed entities to process personal and financial information within Qatar only, with QCB approval before any cloud arrangement. - **UAE health data**: Federal Law 2 of 2019 provides that health data related to services provided inside the state may not be stored or processed outside the UAE except under a health-authority resolution, and it applies including in the free zones. The exemption instrument is Ministerial Decision 51 of 2021; we have not been able to read its operative text, so treat the exception categories as something to confirm with counsel rather than to design around. - **Abu Dhabi health, which is stricter again**: ADHICS V2, issued by the Department of Health in May 2024 and effective from August 2024, requires cloud hosting physically within the UAE with none of the environments, infrastructures or systems outside the country, including backup and disaster recovery, and states no exemption route. A companion control catches encrypted, anonymised and pseudonymised copies, and another prohibits offshore analytics and offshore remote support. For an Abu Dhabi health estate this is the binding constraint, not the federal law. - **UAE banking and payments**: The CBUAE Outsourcing Regulation for banks requires that the Master System of Record, which includes all Confidential Data, is continuously maintained and stored within the UAE. The Consumer Protection Standards extend consumer and transaction data holding to all licensed financial institutions, and the stored value facility and retail payment regulations carry comparable requirements. - **The UAE's three coexisting regimes**: Federal, DIFC and ADGM. The federal PDPL has been in force since January 2022 but its executive regulations have still not been issued, so the compliance clock, the penalty schedule and the adequacy list are not operative, and in June 2026 the Cabinet approved a new AI and Data Authority to absorb the federal Data Office. DIFC entities sit under DP Law 5 of 2020 and Regulation 10; ADGM under its own 2021 regulations. Free-zone entities are not under the federal PDPL. The health data law reaches them anyway. #### A live region in the country does not mean the models run in it This is where most Gulf residency architectures break, and they break quietly, because the region genuinely exists and the platform genuinely runs in it. Azure Qatar Central in Doha is live and hosts Microsoft Foundry. As of September 2026 it serves no Azure-sold models in any deployment type: the region does not appear in Microsoft's model region tables at all, Foundry Models is marked preview there, and the Foundry Agent Service is not offered. A team that reads “Foundry is live in Doha” as “we can run a model in Doha” has mistaken the platform's presence for the model's. AWS shows the adjacent version of the same trap. Bedrock runs in Bahrain me-south-1, and the number of models it serves In-Region from there is zero. Everything reachable from that region is reached over Bedrock's Global routing. The footprint itself needs the same care, and all of it is dated. As of September 2026 AWS runs me-central-1 in the UAE and me-south-1 in Bahrain, with a Saudi region announced and not live. Azure runs UAE North in Dubai and Qatar Central in Doha; UAE Central in Abu Dhabi is a restricted-access disaster-recovery region with no listed products, Qatar Central has no paired region, and Saudi Arabia East is not live. Google runs me-central1 in Doha and me-central2 in Dammam, which is Dammam rather than Riyadh and worth checking against whatever your diagram says, and has no UAE region and none announced. Announced capacity is a plan; a plan is not a control. #### Data at rest is not data in use Even where a model is reachable from a Gulf region, the vendors say plainly that reachability is not residency. Their own documentation is the strongest evidence a buyer has on this, and it is worth reading in the vendors' own terms. AWS separates the two routes in Bedrock. Of In-Region inference it says your requests never leave the Region you specify. Of Global inference it says Bedrock routes your request to a supported commercial Region worldwide, and that data may be processed in any commercial Region. In the Gulf, as of September 2026, exactly two models are In-Region: Amazon Nova Pro and Nova Lite, in me-central-1 only. Every Claude, GPT and Grok model listed in me-central-1 and me-south-1 is Global-only. There is no Middle East Geo tier, the geographies offered being US, EU, Japan and Australia. AWS also notes that input prompts and output results may be stored in the opt-in Regions for abuse detection purposes, which is storage, not merely transit. Microsoft's own deployment-type documentation draws the same line: data stored at rest remains in the designated Azure geography, but Global deployment types may be processed in any Azure region. There is no Middle East Data Zone. In UAE North, the deployment types that are genuinely in-region cover four models as of September 2026: three text-embedding models and whisper, and no chat model. The only in-region route to a frontier chat model there is Regional Provisioned Managed, which requires purchasing committed throughput capacity. Microsoft also describes the order in which deployment types reach a region: Global first, Data Zone next, single-region last, with no guaranteed date for the last of those. Google states it as a warning rather than a footnote: endpoints do not guarantee data residency or in-region ML processing. Its machine-learning processing residency commitments cover thirteen locations, and not one of them is in the Middle East. Assured Workloads offers a Qatar Data Boundary and a KSA Data Boundary, and Google's documentation is explicit that these pin where resources live and do not provide residency controls for data in use or data in transit. Every figure in this section is a snapshot. Vendor model catalogues change frequently and without notice, so re-check each claim against the vendor documentation in the sources below before it turns into an architecture decision, and record the date you checked it. #### What the workable architectures actually are Six paths survive contact with the facts above: five architectures and one procurement gate. None of them is free, and the trade-off in each case belongs in front of whoever signs it off rather than in an appendix. All of the vendor and legal detail here is as of September 2026 and should be re-verified before it is relied on. - **Self-host open-weight models in-region**: SageMaker AI runs in both me-central-1 and me-south-1, and Azure Machine Learning runs in UAE North and Qatar Central. The constraint is accelerators rather than software: on Google Cloud, Dammam lists L4 GPUs only and Doha lists no Vertex AI accelerators at all, and Vertex AI in both is the base platform, without AutoML, model monitoring or agent runtime, which sit in Singapore. You also take on the model operations you were buying a managed service to avoid. - **Buy committed capacity for an in-region managed model**: On Azure UAE North, Regional Provisioned Managed is the route to a frontier chat model that stays in the region. It is a reserved-capacity purchase with a commercial commitment attached, so it belongs in the business case at the start rather than in a late architecture review. - **Use the two models that are actually In-Region**: On AWS in the UAE, Amazon Nova Pro and Nova Lite are the In-Region options. For workloads those models handle well, that is a real architecture. For workloads they do not, it is an honest constraint rather than something to engineer around. - **A sovereign or in-country platform, with the live-versus-announced line held**: Core42's sovereign public cloud on Azure in the UAE is live per its own product pages, scoped to open and confidential or sensitive tiers. Ooredoo's sovereign AI cloud in Qatar has been live since July 2025, per Ooredoo. In Saudi Arabia, AMD reports MI355X capacity live as of 31 August 2026 serving HUMAIN's customers, while the AWS and HUMAIN AI Zone is a build-out its partners plan to 2028, and Stargate UAE's first 200MW is expected rather than confirmed. Buy against what is running, on the provider's own published status rather than the press release. - **Accept cross-border inference where the law permits it**: For personal data outside the sectoral regimes, the transfer route is frequently lawful: Saudi transfers on standard contractual clauses or binding common rules with a transfer risk assessment where required, and Qatari law that positively discourages restricting flow. This is the option most often ruled out by assumption rather than by an instrument. Ruling it in means writing down which instrument applies to which data set, which is the work most residency requirements skip. - **The procurement path, in Saudi Arabia particularly**: Customers billed in Saudi Arabia buy Google Cloud through CNTXT, per Google's own Dammam region documentation, and the Dammam region holds a CST Class C licence assessed by the National Cybersecurity Authority. That is a procurement route rather than an architecture, and discovering it late costs weeks. #### The questions to put to any vendor or consultancy Every question below has a factual answer somebody can look up. A supplier who cannot answer them has not done the work, and a supplier who answers them confidently without a source has done something worse. Take the answers in writing, with the date attached, because all of them expire. - **Which instrument are we complying with, by name?**: Not “Gulf data residency”. The PDPL, ADHICS V2 control CS 1.2, CBUAE Article 6.1, Cloud First Policy P005: the answer should be a named clause, and it determines everything downstream. - **Which model, in which region, on which deployment type?**: All three together, on the vendor's own terms: In-Region or Global on Bedrock; Global, Data Zone, Regional or Regional Provisioned Managed on Azure. A model name and a region name with no deployment type is not an answer. - **Is the region running the model, or only the platform?**: Ask for the line in the vendor's model region table, not a confirmation that the region is live. The two questions have had different answers in Doha and in Bahrain as of September 2026. - **Where is inference performed, and where are prompts and outputs retained?**: Processing location and retention are separate answers. Abuse-detection retention can put prompt text in a region your architecture diagram does not show. - **What happens to backups, disaster recovery, logs and support access?**: The Abu Dhabi health standard names backup and disaster recovery explicitly and prohibits offshore remote support, so an in-country primary with an offshore secondary fails it. Most residency designs are broken by their second copy rather than their first. - **Is this capacity running today, or announced?**: With a date, a source and the name of whoever confirmed it. Sovereign AI announcements in the Gulf run several years ahead of the capacity they describe. - **How would we evidence this to a regulator?**: A region recorded per store, a deployment type recorded per model call, and logs that show what actually ran. If proving it depends on a supplier's recollection, it is not evidence. #### Where we have had to make these decisions We hold end-to-end technical ownership of the live production estate of a government-accredited EdTech in the Middle East. Its database was replatformed onto Postgres in Google Cloud's Doha region, me-central1, to meet Gulf data-residency requirements. That is one region decision on one estate, and it is the basis on which we describe the trade-offs above: not a survey of the market, but the set of choices we had to make, document and defend. The separations on this page are the ones that engagement forced. Where the data sits became a store-by-store decision with a named region against each. Where the models run stayed a separate decision, taken against the vendor documentation set out above rather than against the region's marketing. What we bring to a residency question is architecture and evidence, and the discipline of checking the vendor's own words on the day the decision is made. #### Questions ##### Does the Saudi PDPL require data to stay in Saudi Arabia? No. The PDPL regulates cross-border transfer rather than imposing general data localisation: transfers run on standard contractual clauses, binding common rules or accredited certification, with a transfer risk assessment in defined cases, no general prior approval requirement, and no adequacy list published as at September 2026. The in-country requirements come from elsewhere. The MCIT Cloud First Policy states that all data in both the Government Cloud and the Commercial Governmental Cloud should be located geographically inside the borders of Saudi Arabia, and sector regulators add their own. Identify which of those applies to your data before you design for the strictest reading of all of them. This is a reading of the instruments as at September 2026, not legal advice; verify against the current text and with your own counsel before it becomes a design decision. ##### Can we use Claude or GPT models and still meet Gulf data-residency requirements? As of September 2026, not through the managed routes in Gulf regions as they stand. On AWS Bedrock, every Claude, GPT and Grok model listed in me-central-1 and me-south-1 is Global-only, which by AWS's own definition means the request may be processed in any commercial Region. On Azure, Global deployment types may be processed in any Azure region and there is no Middle East Data Zone. The honest options are: buy committed capacity on Azure UAE North through Regional Provisioned Managed, which keeps a frontier chat model in-region; use Amazon Nova Pro or Nova Lite, the only In-Region models in the Gulf, in me-central-1; self-host open-weight models on in-region compute; or establish that cross-border processing is lawful for that data set and document why. Re-check the model tables before deciding, because they move frequently and without notice. ##### Is a local cloud region enough on its own? No, for two reasons. A region can host the AI platform and none of the models: as of September 2026 Azure Qatar Central runs Microsoft Foundry and serves no Azure-sold models in any deployment type, and AWS Bedrock in Bahrain serves none In-Region. And where a model is reachable, storage location and processing location are separate: Microsoft states that data at rest remains in the designated geography while Global deployment types may be processed in any Azure region, and Google's documentation warns that endpoints do not guarantee data residency or in-region ML processing. A region is a necessary condition for in-country storage and tells you almost nothing about inference. ##### What about the UAE's free zones? Free-zone entities are not under the federal PDPL. DIFC entities sit under Data Protection Law No. 5 of 2020 with Regulation 10, in force since September 2023 and described by DIFC as the region's first data protection rule written specifically for AI processing, introducing concepts including deployers and operators, high risk processing and an autonomous systems officer. ADGM has its own 2021 regulations and no AI-specific regime. Qatar has a parallel split, with QFC entities under separate regulations modelled on the GDPR. None of this exempts you from sectoral rules: the UAE health data law applies including in the free zones. Read the current text from the DIFC Commissioner's office before relying on clause-level detail, because the drafting is what matters here. ##### How long does an answer to this question stay true? Months, not years, on both halves of it. Vendor model catalogues, deployment types and region footprints change on a monthly cadence, and Microsoft's stated order of arrival means single-region deployment types land last and without a guaranteed date. The regulatory side is moving too: the UAE's federal executive regulations have still not been issued, and a new AI and Data Authority was approved by the Cabinet in June 2026. Treat any residency statement, including this page, as carrying the date it was checked, and re-verify against the primary sources before a design or a contract depends on it. #### Sources [1] SDAIA, Personal Data Protection Law and regulations: https://sdaia.gov.sa/en/SDAIA/about/Pages/RegulationsAndPolicies.aspx [2] Saudi MCIT, Cloud First Policy: https://www.mcit.gov.sa/sites/default/files/cloud_policy_en.pdf [3] UAE Federal Decree-Law No. 45 of 2021 (Personal Data Protection): https://uaelegislation.gov.ae/en/legislations/1972 [4] UAE Federal Law No. 2 of 2019 (ICT in Health Fields): https://uaelegislation.gov.ae/en/legislations/1209 [5] CBUAE, Outsourcing Regulation for Banks (C 14/2021): https://rulebook.centralbank.ae/en/rulebook/outsourcing-regulation-banks [6] Department of Health Abu Dhabi, ADHICS V2 standard: https://www.doh.gov.ae/-/media/Feature/Resources/Standards/ADHICS-v2-standard.ashx [7] Qatar Law No. 13 of 2016 (PDPPL), official English translation, Al-Meezan legal portal: https://www.almeezan.qa/EnglishLaws//132016.pdf [8] Qatar Central Bank, Cloud Computing Regulation (2024): https://www.qcb.gov.qa/Documents/InformationSecurity/Cloud%20Computing%20Regulation.pdf [9] Hukoomi, Qatar government portal, carrier of the MCIT Cloud First Policy (P005): https://hukoomi.gov.qa/ [10] DIFC Commissioner of Data Protection, DP Law No. 5 of 2020 and Regulation 10: https://www.difc.com/business/registrars-and-commissioners/commissioner-of-data-protection [11] AWS, Bedrock model region compatibility: https://docs.aws.amazon.com/bedrock/latest/userguide/models-region-compatibility.html [12] Microsoft, Azure AI Foundry deployment types and data residency: https://learn.microsoft.com/azure/ai-foundry/foundry-models/concepts/deployment-types [13] Microsoft, Azure regions list: https://learn.microsoft.com/azure/reliability/regions-list [14] Google Cloud, regions and zones: https://cloud.google.com/compute/docs/regions-zones [15] Google Cloud, generative AI locations and data residency: https://docs.cloud.google.com/gemini-enterprise-agent-platform/resources/locations [16] Google Cloud, Dammam region access and CNTXT purchasing: https://docs.cloud.google.com/docs/dammam-region-access Last reviewed 1 September 2026. ## Credential register (https://www.1aym.com/claude-certified-architect) The credential register: both Claude certifications held by 1AYM’s founder and Principal Architect, rendered verbatim from Credly’s public record. ## OpenAI Select Partner (https://www.1aym.com/openai-select-partner) Company status: 1AYM is an OpenAI implementation partner in the UK and an OpenAI Select Partner, awarded after a partner agreement, a compliance review and a technical assessment. Not a certification from OpenAI. 1AYM holds no partner status with Anthropic. Model-selective production systems, client environments, five named engagements. ## Every route - https://www.1aym.com/ · 1AYM homepage - https://www.1aym.com/about · About 1AYM - https://www.1aym.com/capabilities · Services - https://www.1aym.com/capabilities/executive-ai-discovery · AI Opportunity & Feasibility Sprint - https://www.1aym.com/capabilities/ai-harness-platform-engineering · AI harness & platform engineering - https://www.1aym.com/capabilities/production-ai-systems · Production AI systems - https://www.1aym.com/capabilities/agentic-workflow-design · Agentic workflow design - https://www.1aym.com/capabilities/data-platform-ai-enablement · Data platform & AI enablement - https://www.1aym.com/capabilities/enterprise-integrations-automation · Enterprise integrations & automation - https://www.1aym.com/capabilities/fractional-ai-platform-architect · Fractional AI platform architect - https://www.1aym.com/capabilities/embedded-engineers-on-contract · Embedded Engineers on Contract - https://www.1aym.com/capabilities/ai-governance-implementation · AI governance implementation - https://www.1aym.com/work · Engagement files - https://www.1aym.com/work/finance-data-access-connector · Self-serve finance answers, without breaking permissions - https://www.1aym.com/work/education-ai-enablement · An IELTS speaking assessment built as a measuring instrument - https://www.1aym.com/work/ai-platform-enablement · AI, data and automation enablement across a global agency - https://www.1aym.com/work/identity-provisioning-at-scale · Identity provisioning for 1,000+ users across 600+ groups - https://www.1aym.com/work/clinical-trials-platform-rescue · A regulated product that was not going to ship, shipped. - https://www.1aym.com/claude-certified-architect · Claude Certified Architect - https://www.1aym.com/openai-select-partner · OpenAI Select Partner - https://www.1aym.com/what-is-an-ai-harness · What is an AI harness? - https://www.1aym.com/agent-proposes-verifier-gates · Agent proposes, verifier gates - https://www.1aym.com/how-to-choose-an-enterprise-ai-consultancy · How to choose an enterprise AI consultancy in the UK - https://www.1aym.com/codex-and-claude-code-enterprise-rollout · How to roll out Codex and Claude Code across an engineering organisation - https://www.1aym.com/gulf-ai-data-residency · Gulf AI data residency: the architecture questions buyers should ask - https://www.1aym.com/privacy · Privacy notice - https://www.1aym.com/llms.txt · llms.txt - https://www.1aym.com/llms-full.txt · llms-full.txt ## Contact - Email: tayyeb@1aym.com - Book a 30-minute call: https://cal.com/tayyeb-mahmud-jxedia - Sitemap: https://www.1aym.com/sitemap.xml - Short form of this file: https://www.1aym.com/llms.txt Clients are referred to by descriptor in this file and on every route, with two exceptions: the engagement file whose published source names its client, and the homepage logo bar, which shows client marks. Descriptors are used everywhere else.