# 1AYM full text > 1AYM is an enterprise AI and platform consultancy and an OpenAI Select Partner. We take organisations from AI ambition to governed production systems their people use every day. Generated from the same content modules that render https://www.1aym.com, so nothing here can describe a page that does not exist. British English throughout. Last reviewed 22 August 2026. Clients are not named in this file. Where a figure came from our own measurement rather than a client publication, the source line says so, exactly as it does on the page. ## Key Facts - Type: enterprise AI and platform consultancy, headquartered in Stoke-on-Trent, United Kingdom. Language: en-GB. - Delivery: UK, the Gulf and the US. Registered in England and Wales. - Company status: OpenAI Select Partner, awarded after a partner agreement, compliance review and technical assessment. 1AYM is not an Anthropic or Claude partner, and "OpenAI certified" is not a thing 1AYM is, or claims. - Certifications: Claude Certified Architect – Professional (Anthropic, 23 July 2026) and Claude Certified Associate – Foundations (Anthropic, 21 July 2026), held by the founder and Principal Architect, independently verifiable on Credly. - Compliance: we work to UK GDPR and the Data Protection Act 2018. We hold no ISO or SOC certification. - Clients: referred to by descriptor in this file and on every route, with two exceptions. One engagement file names its client, because the published source it links carries the name. The homepage logo bar shows two client marks. - Capability areas: 8. Named engagements: 5. Published engagement files: 5. Client environments in daily production use: 2. ## Who we are The work is executive discovery, architecture, hands-on engineering, governance, rollout and measurement. Where we work. We are headquartered in the UK and deliver across the UK, the Gulf and the US. One client environment runs on Postgres in Google Cloud’s Doha region (me-central1) to meet Gulf data-residency requirements, and another reached 91% adoption in North America, a figure published in the Anthropic customer case study. Who does the work. Every 1AYM engineer is either certified on the platforms we build on or has shipped inside a top-tier engineering organisation: Meta, Spotify, UBS, Starling Bank, S&P Global, Sky. We scale the team to the contract, and the people who write the strategy write the code. The senior engineers a client meets first: a Principal Architect (10+ years; AI adoption, agent systems, Claude and OpenAI platform architecture; holds both Anthropic Claude certifications; in production at both client environments below); a Principal Engineer (14 years, London; data platform, context and memory architecture; shipped inside Meta, Spotify, Viasat/Inmarsat); a Principal Data Engineer (17 years, London; pipeline orchestration, CI/CD, data infrastructure; shipped inside Starling Bank, UBS, S&P Global, Sky, Société Générale). The Principal Architect is Tayyeb Mahmud, who founded the practice. 1AYM delivers the work rather than any one person; other colleagues are named only with their permission, and clients meet the engineers before signing. Meta, Spotify, Viasat/Inmarsat, Starling Bank, UBS, S&P Global, Sky and Société Générale are former employers of our engineers, not clients of 1AYM. Certifications and status are two different things and 1AYM keeps them apart. The Claude Certified Architect – Professional (Anthropic, issued 23 July 2026) and Claude Certified Associate – Foundations (Anthropic, issued 21 July 2026) are held by 1AYM’s founder and Principal Architect, and are independently verifiable on Credly, which carries the holder’s public record. 1AYM’s OpenAI Select Partner status is a company status, held by the practice rather than by any individual, and was awarded after a partner agreement, compliance review and technical assessment. "OpenAI certified" is not a thing 1AYM is, or claims: the correct wording is "OpenAI Select Partner and Claude certified". 1AYM is not an Anthropic or Claude partner. No vendor bias. Model choice at 1AYM comes from production evidence across OpenAI, Anthropic, Gemini and open-source, not from a vendor deal. We are model-selective across OpenAI, Anthropic, Gemini and open-source components, chosen per workload rather than by vendor loyalty. In production today: OpenAI audio models for speech assessment with a Gemini fallback, Claude for enterprise platform and agent estates, Codex and Claude Code for engineering work. We work to UK GDPR and the Data Protection Act 2018, and hold no ISO or SOC certification. ## Homepage (https://www.1aym.com/) **Enterprise AI your people use every day.** The people who write the strategy write the code. Every 1AYM engineer is certified or has shipped inside Silicon Valley, banking and government-scale systems, and we scale the team to the contract. OpenAI Select Partner and Claude certified. We run OpenAI, Anthropic and Gemini models in production, and open-source where it fits, so the choice for your use case comes from evidence rather than a vendor deal. We are headquartered in the UK and deliver across the UK, the Gulf and the US. ### How it works Three stages, in this order. Each one ends with something written down and handed over, so you can stop after any of them and still own what you paid for. #### Discover Typically two to four weeks with your sponsor and the people who do the job today. We agree what the outcome is worth, what stands in the way, and how it will be judged. You leave with a ranked list of what to build, the architecture for it, and an investment case with real numbers in it. #### Ship a working slice One real workflow, end to end, in your systems and against your data. A demonstration that skips the hard integration proves nothing, so we do not build one. You leave with working software in your environment, and evidence of what it does. #### Run it, governed Access rules, approval gates, monitoring and training, then a staged rollout. Your team runs it. We measure adoption, cycle time and cost, so the change is reported rather than asserted. You leave with a running system, a team that can maintain it, and numbers for the board. ### What you can buy Five named engagements. Each one has a buyer, a stated output or a clear commercial shape, and paperwork you can put in front of procurement without a rewrite. #### AI Opportunity & Feasibility Sprint For C-suite and transformation leads. Fixed scope · typically 2–4 weeks. Find out which two or three workflows are worth building, what each is worth and what it would cost, before you commit a budget to any of them. You get: - A ranked shortlist of use cases, with the value and the risk against each one - The architecture, the data and integration constraints, and which models fit - A delivery roadmap and an investment case your finance team can sign off #### Enterprise AI Platform & Agentic Workflows For CIO, CTO, CDO, operations and finance. Fixed-scope architecture, then a production build. Turn scattered pilots into one governed place for AI to run, then put real workflows on it. Access rules, checks and human approval are built in, so the second workflow does not start from scratch the way the first one did. You get: - Platform architecture: who can see what, and the context an AI system is allowed to use - Agents that act on real systems, checked against fixed rules and approved by a human before anything lands - Testing, monitoring and an operating model your team owns #### AI Engineering Transformation (Codex & Claude Code) For engineering leaders. Typically 6–12 weeks · enablement, guardrails and CI. Your engineers are already using AI tools, with or without a policy. This puts review, guardrails and CI around that, so the code it writes reaches production the same way every other change does. You get: - Codex and Claude Code workflows mapped to how your teams actually ship - Repository guardrails, CI gates and multi-agent delivery patterns - Training, documentation and adoption figures you can report on #### Fractional AI Platform Architect For scale-ups and enterprises. Retained · typically 2–3 days a week · one statement of work. Senior architecture and delivery ownership without hiring a permanent AI platform leader. We hold the technical decisions, and the standard the work is judged against. You get: - Architecture ownership, with the model and vendor decisions written down and dated - Delivery governance and technical review of work already in flight - Team enablement, so the capability stays when the engagement ends #### Embedded Engineers on Contract For programmes that need senior capacity now. Contract · from three months · priced per day. A senior engineer, or a team, inside your programme on contract. You direct the work day to day, and we hold the standard every 1AYM engineer is hired against, plus the cover behind them. You get: - A named senior engineer or a team inside your programme, typically within weeks - The 1AYM standard, and cover across the stream so no role hangs off one person - A clean handover, or conversion to a fixed-scope build once the work is scoped ### Production proof Two client environments depend on this work every working day. Every figure is either the client's published case-study number or our own count or measurement on the systems we run, and each one says which. #### A PE-backed marketing agency We lead the architecture of the enterprise AI platform underneath the agency's organisation-wide AI programme. That platform is featured in a published Anthropic customer case study, which describes the platform. What we built: - The publishing pipeline and release controls that let non-technical staff author AI skills and Claude run them - An OAuth-secured organisational-context MCP server with default-deny permissions and a full audit trail - Automated access provisioning from the HR system, and reporting on what AI actually costs - A re-architecture of the skills estate to the Anthropic SDK's structure, with reference files loaded on demand and MCP servers that inject targeted context in place of whole files, which by our own measurement cut the average cost per session by 60% - ~1,000 · Employees reached (source: Anthropic customer case study) - ~400 · Skills authored in four weeks (source: Anthropic customer case study) - 91% · Adoption in North America (source: Anthropic customer case study) #### A government-backed EdTech We hold end-to-end technical ownership of the live production estate of a Middle East education company: the exam platform, the APIs and the data, replatformed onto Postgres 16 in Google Cloud's Doha region to meet Gulf data-residency requirements. Its IELTS speaking assessment is not one model call wrapped in a product. It is a measuring instrument, built to Ministry of Education evidence standards and in daily use by the client's own team. What we built: - Whole-interview speaking assessment: the official IELTS band descriptors applied in ordinary code, so every band can be reproduced and inspected line by line - Five layers of evidence, from deterministic acoustic measurement to a phoneme-level pronunciation model, with the AI never doing the arithmetic - Confidence routing to a human examiner, and accent fairness and agreement with examiners as release gates the build has to pass before it ships - 3,100+ · Students on the platform (source: Production database, 21 August 2026) - 1,000+ · Added in the last 90 days (source: Production database, 21 August 2026) - 1,400+ · Automated tests, from ~300 (source: Our count, in CI) #### Measured outcomes, client unnamed - £2–4M · Forecast annual saving (source: Our model of token consumption, reviewed by the client's finance team) - 46% · Faster first response (source: Our measurement · client unnamed) - In the meeting · Governed finance answer, from ~2 hours (source: Our measurement · client unnamed) - 82% · Less manual triage (source: Our measurement · client unnamed) ### 1AYM vs the alternatives There are four ways to buy this kind of work. They differ in who writes the code, how long it takes to reach production, and what you still own once the invoice is paid. | | 1AYM | Big-4 / systems integrator | Staff augmentation | Strategy-only advisory | | --- | --- | --- | --- | --- | | Who does the work | Certified engineers, or engineers who shipped inside Meta, Spotify, UBS and Starling Bank. Senior, hands-on, no bench of juniors. | A partner sells it; a large mixed-seniority team delivers it. | People you brief, manage and direct yourself. | Advisers who hand the build to somebody else. | | Time to first production slice | Typically weeks. One real workflow, end to end, in your environment. | Months, after a discovery phase and a mobilisation phase. | As fast as the plan you give them, and no faster. | None. The output is a document. | | Model choice | From production evidence across OpenAI, Anthropic, Gemini and open-source, not from a vendor deal. | Usually shaped by the vendor alliance and reseller economics the firm already has. | Whatever the people you hired happen to know. | A recommendation, with nothing in production behind it. | | Governance | Built as platform controls: default-deny access, permission-aware context, deterministic checks, human approval, evaluation gates in CI. | A governance workstream, usually a policy set and a steering committee. | Yours to define, enforce and audit. | A framework for somebody else to implement. | | What you are left with | Running systems, the architecture written down, and a team that can extend both. | A system your team did not build, and a change-request route into it. | Whatever got built. The knowledge leaves when they do. | A roadmap, and the build still to buy. | | Commercial shape | Fixed-scope statements of work, a retained architecture engagement, or embedded engineers from three months. One supplier, one contract. | Multi-phase programme against a rate card. | Priced per person, per day, open-ended. | Fixed fee for the report. | ### Six things we do not trade None of them is a methodology. They are the constraints we hold on every engagement, and the reason the work is still running after we leave. - **Certified or proven**: Every engineer is certified on the platforms we build on, or has shipped inside a top-tier engineering organisation. - **Fixed scope**: Defined outputs, a defined price and acceptance criteria in the statement of work, or a clear contract. - **Governed from day one**: Access control, approval gates and evaluation in CI are part of the build, not a later project. - **Your people keep the keys**: Architecture, code and operating model are documented and handed over. Nothing depends on us staying. - **No vendor bias**: Chosen per workload across OpenAI, Anthropic, Gemini and open-source, on production evidence rather than vendor loyalty or a vendor deal. - **Overlapping coverage**: The engineers on your work cover each other's streams, so one person's calendar is never the critical path. ### Certified, or proven at the top Every engineer here is either certified on the platforms we build on, or has shipped inside a top-tier engineering organisation: Meta, Spotify, UBS, Starling Bank, S&P Global, Sky. We scale the team to the contract. - **Tayyeb Mahmud · Founder and Principal Architect** (10+ years) · AI adoption, agent systems, and Claude and OpenAI platform architecture. Sits with the sponsor, sets the standard, and still writes the code. In production: A PE-backed marketing agency · a government-backed EdTech. Holds both Anthropic Claude certifications, Architect – Professional and Associate – Foundations, verifiable on Credly. - **Principal Engineer** (14 years · London) · Data platform work, and the context and memory architecture that decides what an AI system is allowed to see, keep and recall. Shipped inside: Meta · Spotify · Viasat/Inmarsat. - **Principal Data Engineer** (17 years · London) · Pipeline orchestration, CI/CD and data infrastructure: the layer most AI programmes are actually blocked on once the model choice is settled. Shipped inside: Starling Bank · UBS · S&P Global · Sky · Société Générale. We name colleagues only with their permission, and we do not name clients. 1AYM delivers the work, and the role, the years and the organisations on each card are the standard being met. You meet the engineers before you sign anything. 1AYM is an enterprise AI and platform consultancy and an OpenAI Select Partner, founded by Tayyeb Mahmud. Headquartered in the UK, we deliver across the UK, the Gulf and the US. We take organisations from AI ambition to governed production systems their people use every day. Two client environments depend on that work today. ### Frequently asked questions #### Who actually does the work? Certified engineers, or engineers who have shipped inside Meta, Spotify, UBS, Starling Bank, S&P Global and Sky, and the same people who scoped the work. There are no juniors on it. Our principal architect leads the architecture and holds the Claude Certified Architect – Professional and Claude Certified Associate – Foundations certifications, both verifiable on Credly. A principal engineer of fourteen years takes the data platform and the context and memory layer; a principal data engineer of seventeen takes pipeline orchestration, CI/CD and data infrastructure. They cover each other's streams, and we scale the team to the contract. 1AYM delivers the work rather than any one person, we name colleagues only with their permission, and you meet the engineers before you sign anything. #### How do engagements start? With a 30-minute call to work out whether your problem is one we should take. If it is, the first paid piece is fixed-scope: usually a two-to-four week opportunity and feasibility sprint, or a first production slice against one real workflow. We will say plainly when the answer is that you do not need us. Expect weeks rather than quarters to a first working slice, because the delay on these projects is almost never the model. It is access, data quality and the approval path, so we start those on day one. #### How do you contract? Three shapes. Most work is a fixed-scope statement of work with defined outputs and a defined price: a sprint, an architecture and build, or an enablement programme. The second is a retained architecture engagement, typically two to three days a week, under a single statement of work. The third is embedded engineers, a senior engineer or a team inside your own programme on a contract from three months, where the constraint is capacity rather than a scoped deliverable. In every case it is one supplier and one contract. A statement of work carries the outputs, the price and the acceptance criteria; an embedded contract carries the roles, the minimum term and the review cadence. Where a programme needs more engineers we scale the team to the contract, and every one of them is held to the same standard. #### What does “production” mean to you? A system real people depend on in their working day. That means identity and access control, monitoring, deterministic checks where the output matters, a human approval path for anything high-risk, and a named owner inside your organisation. Something that runs on an engineer's laptop is a prototype. Two client environments meet that definition today. #### How do you handle governance and data? Governance is built as controls rather than written as a policy: least-privilege and default-deny access, permission-aware context so an AI system sees only what the user is already entitled to see, audit trails that record refusals as well as answers, deterministic validation, human approval sized to risk, and evaluation gates in CI. We work to UK GDPR and the Data Protection Act 2018, plus whatever your sector adds. Where data must stay in a region we architect for it: one production estate runs on Postgres in Google Cloud's Doha region. We hold no ISO or SOC certification. #### Which models do you use? Whichever one the workload argues for. The choice comes from what we run in production across OpenAI, Anthropic, Gemini and open-source, not from a vendor deal. We are model-selective across OpenAI, Anthropic, Gemini and open-source components, chosen per workload rather than by vendor loyalty. In production today: OpenAI audio models for speech assessment with a Gemini fallback, Claude for enterprise platform and agent estates, Codex and Claude Code for engineering work. Where a model is a poor fit for a job we will say so and price the alternative. ## Services (https://www.1aym.com/capabilities) Index of the eight capability areas. The nav calls them Services; the route is /capabilities. ### Executive AI discovery & product translation (https://www.1aym.com/capabilities/executive-ai-discovery) Reference: C-01. Question this page answers: How do we turn board-level AI ambition into a buildable plan? Sold as: AI Opportunity & Feasibility Sprint. Fixed scope · typically 2–4 weeks. Executive AI discovery turns an ambition into a plan a CFO can fund and an engineer can start on Monday. We run the leadership sessions, map where the value sits, test feasibility against your data and systems, settle build versus buy, and write the roadmap, the investment case and the governance conditions. Programmes rarely stall because the technology is hard. They stall because nobody turned the ambition into something specific, sequenced and fundable, and the cheapest place to find that out is before the budget is committed. #### The problem this solves A board asks for an AI strategy. What comes back is either a slide deck with no engineering in it, or an engineering plan with no business case in it. Both stall, for the same reason: the person writing them could only see one half of the problem. Discovery closes that gap. It produces a plan that a CFO can fund and an engineer can start on Monday, because the same person wrote both halves and can defend the trade-offs between them. #### How the work runs It is deliberately short. The aim is a decision, not a documentation exercise. - **Leadership sessions**: What the business is actually trying to change, commercially rather than technically. Where the pressure is coming from, and what success would look like to the people funding it. - **Workflow analysis**: How the work is done today, including the spreadsheets and manual steps nobody documents. This is usually where the real opportunities are hiding. - **Data and systems feasibility**: What data exists, what state it is in, who is allowed to see it, and which systems would have to be touched. Most AI plans die here, and it is cheaper to find out early. - **Opportunity map**: Candidate use cases scored on value, feasibility and risk, so the sequence is defensible rather than whichever idea had the loudest sponsor. - **Governance review**: What has to be true for legal, security, and risk to sign off, established before anything is built rather than discovered at the end. #### Why it comes first Almost every expensive AI failure traces back to skipping this. A team builds the use case that was easiest to describe, rather than the one that was most valuable, and finds out at rollout that the data was not accessible or the process owner was never consulted. Discovery is the cheapest step in the programme and the one that determines whether the rest of the money is well spent. #### What you get - Opportunity map with value, feasibility and risk scoring - Current-state workflow analysis - Technical architecture for the recommended direction - Sequenced delivery roadmap - Risk and governance notes - Build-versus-buy recommendation #### Start here if - Leadership has committed to AI but nobody has converted it into a plan - Several teams are running disconnected experiments - A business case is needed before budget is released - Previous AI work stalled and no one is certain why #### Questions ##### How long does discovery take? It is scoped as a short, focused engagement rather than an open-ended consulting phase: long enough to interview the people who own the work and assess the data honestly, short enough that it does not become the project. The output is a decision and a roadmap, not a research programme. ##### Do you need access to our data to do this? It helps a lot. Feasibility claims made without seeing the data are guesses. Where access is impossible in the timeframe, the assessment states its assumptions so they can be tested before build. Last reviewed 27 July 2026. ### AI harness & proprietary platform engineering (https://www.1aym.com/capabilities/ai-harness-platform-engineering) Reference: C-02. Question this page answers: How do we build a reusable internal AI platform? Sold as: AI Engineering Transformation (Codex & Claude Code). Enablement, guardrails and CI · typically 6–12 weeks. An AI harness is the platform layer that connects models, tools, workflows, business context and permissions into one operating layer, so the organisation ends up with reusable capability rather than a shelf of one-off agents. The test of whether you need one is simple: if your third AI project shares nothing with the first two, you are paying for the same plumbing three times. It is extracted from working software as real use cases land, never designed in advance of them. #### What gets built A harness is the set of layers every AI use case in the organisation draws on. Built once, governed centrally, and inherited by everything downstream. - **Shared operating layer**: The substrate each use case is built on, so the second and third cost a fraction of the first. - **Agent and skill framework**: Reusable capabilities that compose, rather than a separate bespoke agent per department. - **Tool and model routing**: Which model handles which step, which tools it may call, and how the system degrades when one is slow or unavailable. - **Governance and permissions**: What each agent may read and write, on whose behalf, under which approvals, enforced by the platform rather than reimplemented per team. - **Monitoring and feedback loops**: What ran, what it produced, whether it was right, and a route for corrections to improve the system. #### Built alongside use cases, not before them Platform projects that precede real usage produce abstractions the business never needed. The route that works is to deliver one production use case properly, then extract the shared layers as the second and third are built. That sequencing keeps the platform honest: every layer in it exists because a real workflow demanded it. #### What it is not It is not an agent framework: that is a library you might use inside a harness. It is not a model subscription, which gives individuals a better tool but gives the organisation no capability. And it is not a product you install, because the routing, context and permission model encode decisions specific to your business. #### What you get - Shared operating layer and platform architecture - Agent and skill framework - Tool and model routing - Governance and permission model - Monitoring, evaluation and feedback loops - Documentation for the teams building on top #### Start here if - A third AI use case is being scoped and the first two shared nothing - Each new use case costs as much as the last - No consistent answer to governance questions across AI projects - Teams are reimplementing the same integrations and permissions #### Questions ##### Does a harness lock us into one model provider? The opposite, when it is built correctly. Routing is one of the layers the harness owns, which makes model choice a configuration decision rather than an architectural one. Mixing or swapping providers per task is a property you gain from having a harness, not something you trade away. ##### Can we build this on top of an existing agent framework? Usually, and often you should. What a framework cannot supply is your business context, permission model and evaluation suites: the part that makes the platform yours. Last reviewed 27 July 2026. ### Production AI systems (https://www.1aym.com/capabilities/production-ai-systems) Reference: C-03. Question this page answers: How do we get an AI system into production and keep it there? Sold as: Enterprise AI Platform & Agentic Workflows. Fixed-scope architecture plus production build. A demonstration has one happy path, a clean input and somebody watching. A production system has every input your business can generate, nobody watching, and a consequence attached to being wrong. We build the second kind: internal copilots, workflow agents, voice agents and the middleware underneath them, designed around who is allowed to see what, what happens when a model fails, what gets logged, and where a human signs off. That is the part that decides whether it is still running in a year. #### The gap between a demo and a system A demo has one happy path, a curated input, and a human watching. A production system has every input the business can generate, no one watching, and a consequence attached to being wrong. Closing that gap is most of the engineering. It means deciding what happens when the model is unavailable, when the input is malformed, when the output fails validation, when a downstream system rejects a write, and when someone needs to know six months later why a particular decision was made. #### What the design actually covers - **Data access and permissions**: What the system can read and write, acting on behalf of whom, and how that is enforced rather than assumed. - **Failure modes**: What happens on timeout, malformed output, rate limit, partial write, and downstream rejection, each one designed rather than discovered. - **Evaluation**: Regression suites that run when prompts or models change, so a quality drop is caught before users find it. - **Human review**: Where a person sits in the loop, what they are shown, and how their corrections feed back. - **Auditability**: A record of what ran, on which inputs, producing which output: the thing that makes the system answerable afterwards. #### Safe by default The engineering defaults matter more than the model choice: idempotent operations so a retry cannot double-write, dry-run modes so a change can be inspected before it lands, audit logs as a by-product rather than an afterthought, and kill switches that work. These are unglamorous and they are the difference between automation that is clever and automation that is safe to depend on. #### What this looks like in production Three figures from systems we run, each one published in full on the engagement file it links to. - ~1,170 automated tests cover the speaking pipeline on one production estate, among them a regression suite where disabling any one guardrail breaks specific frozen cases. (source: Our count, in CI; https://www.1aym.com/work/education-ai-enablement) - ~300 to 1,400+ the automated test estate on that same platform, over the course of the engagement. (source: Our count, in CI; https://www.1aym.com/work/education-ai-enablement) - 60% lower average cost per session on another client's AI platform after we re-architected its skills estate. (source: Our measurement on the platform; https://www.1aym.com/work/ai-platform-enablement) #### What you get - Working system running against production data - Evaluation and regression suites - Human review paths and escalation routes - Audit logging and observability - Runbook and handover documentation #### Start here if - A prototype works but nobody will let it near production - An AI feature shipped and quietly degraded - There is no way to tell whether output quality has changed - The system needs to write to a system of record #### Questions ##### How do you evaluate an AI system that has no single right answer? By separating the parts that do from the parts that do not. Structure, schema, and factual grounding can be checked deterministically. Genuinely subjective quality is scored against a fixed set of graded examples, so the question becomes whether today's output is worse than last week's rather than whether it is perfect. ##### What happens when the model provider changes the model underneath us? This is exactly what the regression suite exists for. Without one, a silent provider-side change is discovered by users. With one, it is caught on the next run and the routing layer can pin or reroute while the difference is assessed. Last reviewed 22 August 2026. ### Agentic workflow design (https://www.1aym.com/capabilities/agentic-workflow-design) Reference: C-04. Question this page answers: How do we make AI automation dependable enough for finance or compliance? A model that is nearly always right sounds excellent until it runs against your ledger all day. Nearly right is wrong, repeatedly, and delivered with exactly the same confidence as the correct answers. So the model is never the last thing to touch a decision. The agent proposes; deterministic rules, schemas, tests and a human where the risk warrants it decide what proceeds. Finance, data migration and compliance work is built this way because an auditor, not a user, finds the errors. #### Why accuracy is the wrong measure A model that is right 95% of the time produces fifty wrong entries a day at a thousand runs, delivered with the same confidence as the correct ones. The useful question is not how often the system is right, but whether anything notices when it is not. The pattern that answers it is set out in full on the concept page, with worked detail on what a verifier actually checks. - **The agent proposes**: The judgment across messy inputs. - **The verifier gates**: Deterministic checks, none of them a model. - **Pass proceeds, fail holds**: A failure routes to a person with the reason attached. #### Where it earns its keep Finance-critical automation, data migration validation, knowledge extraction, semantic-layer query validation, report generation, operational triage, and internal tooling that writes to systems of record. It is unnecessary where the output is read by the person who asked for it before anything happens: a drafting assistant does not need a gate, because the reader is one. #### What you get - Workflow design with explicit gate criteria - Deterministic validator suite - Human review routing and exception handling - Audit trail generated as a by-product of the gates - Test coverage for the verification layer #### Start here if - Automation touches money, identity, or a regulated process - A previous automation shipped errors nobody caught - Reviewers are rubber-stamping a queue that is mostly correct - The process must be explainable to audit or risk #### Questions ##### What do you need from us before gate design can start? The workflow, and the cost of a silent error in it. Those two facts decide how much of the check can be deterministic and where a person has to stand. Everything else follows from them. ##### Does this work on a process we have already automated? Usually, and it is the more common starting point. An automation that already runs has a history of the errors it produced, which is the best specification anyone could write for what the verifier has to catch. Last reviewed 22 August 2026. ### Data platform & AI enablement (https://www.1aym.com/capabilities/data-platform-ai-enablement) Reference: C-05. Question this page answers: How do we connect AI to data the business actually trusts? The expensive failure in enterprise AI is a plausible answer that disagrees with the official number by a few percent, for reasons nobody can trace. Two of those and the finance team stops trusting the tool, whatever the model behind it. We build the layer that prevents it: warehouse integrations, definitions agreed once in a semantic layer, data contracts and reconciliation, so an AI answer lands on the same figures as the board pack. The trust is built in the data layer, underneath the model. #### The trust problem The most common failure in enterprise AI is not a wrong answer. It is a plausible answer that disagrees with the official number by four percent, for a reason nobody can trace. Once that happens twice, the system is dead. People stop using it, because checking its output costs more than doing the work manually. Trust is the actual product, and it is built at the data layer rather than the model layer. #### What the work involves - **Warehouse integration**: Making Snowflake, BigQuery or the equivalent reachable by AI systems under the same access controls that govern everyone else. - **Semantic-layer enablement**: Exposing the business's own definitions (what counts as an active customer, a booked deal, a period), so a model inherits them rather than inventing its own. - **Data contracts**: Explicit agreements about shape, meaning and freshness, so an upstream schema change surfaces as a failed contract rather than a silently wrong answer. - **Reconciliation workflows**: Checks that an AI-produced figure agrees with the system of record, run automatically rather than trusted. - **AI-accessible reporting**: Reporting surfaces designed to be queried by a system, not only rendered for a human. #### Why semantics matter more than retrieval Most retrieval problems in enterprise AI are really definition problems. The model finds the right table and still produces the wrong number, because the business has three definitions of revenue and the model picked one. Fixing that at the semantic layer fixes it for every use case at once. Fixing it in prompts fixes it in one place until someone writes a new prompt. #### What you get - Warehouse and semantic-layer integration - Data contracts covering shape, meaning and freshness - Reconciliation checks against systems of record - AI-accessible reporting surfaces - Documentation of definitions the AI layer inherits #### Start here if - AI answers disagree with the official numbers - The business has multiple definitions of the same metric - An upstream schema change silently broke a downstream answer - Analysts are re-checking everything the AI produces #### Questions ##### Do we need a semantic layer before we can use AI on our data? Not before, but you will end up building one. Without shared definitions, each use case encodes its own interpretation in prompts and queries, and they drift apart. Formalising the definitions once is cheaper than reconciling five accidental versions of them later. ##### Is this the same as building a RAG system? No. Retrieval decides which documents or rows the model sees. This decides what the numbers in them mean and whether they agree with the system of record. A retrieval system on top of undefined data retrieves the wrong number faster. Last reviewed 27 July 2026. ### Enterprise integrations & automation (https://www.1aym.com/capabilities/enterprise-integrations-automation) Reference: C-06. Question this page answers: How do we put AI inside the tools our teams already use? AI in a separate tab does not change how work gets done. Nobody switches to a chat window for a task their existing system already half handles, so the value appears when it shows up in the ticket, the ledger or the record. We build it into the systems you already run: NetSuite, Xero, Slack, Notion, identity and HR feeds. Then we make it safe to depend on: dry-run modes, idempotent writes, rollback paths, audit logs, and never more access than any person in the company holds. #### Integration is the product AI that lives in a separate tab does not change how work gets done. People will not context-switch into a chat window to do a task their existing system already half-handles. The value appears when the capability shows up inside the workflow: in the ticket, the ledger, the channel, the record. That makes integration the deliverable rather than the plumbing behind it. #### The unglamorous parts that decide whether it survives - **Idempotency**: A retry must not double-write. Systems fail midway more often than they fail cleanly, and a non-idempotent sync turns a transient error into a data problem. - **Dry-run modes**: Every destructive operation should be inspectable before it lands, so a change can be reviewed rather than trusted. - **Rollback paths**: A defined way back from a bad run, decided before the bad run rather than during it. - **Audit logs**: What changed, when, on whose behalf, and why: a by-product of running rather than a compliance project afterwards. - **Rate limits and backoff**: Enterprise APIs throttle. Handling that properly is the difference between a sync that completes and one that half-completes nightly. #### Identity and permissions Anything that reads or writes on a user's behalf inherits that user's permissions, which means identity is part of the integration rather than a layer above it. SCIM provisioning, role mapping, and de-provisioning all have to work, including the unhappy paths. Getting this wrong is how an automation ends up with more access than any human in the organisation. #### What you get - Integration pipelines with idempotent writes - Dry-run and rollback tooling - Audit logging and run observability - Identity and permission mapping, including de-provisioning - Scheduling with predictable, monitored runs #### Start here if - AI capability exists but nobody uses it because it is in another tool - A sync fails partway and leaves inconsistent state - No audit trail for automated writes - Access granted to an automation that nobody reviews #### Questions ##### Can you work with our existing integration platform? Usually yes. Where a tool like a workflow automation platform already handles the orchestration adequately, the sensible move is to use it and add the missing guarantees around it. Custom middleware is worth building when the guarantees matter more than the convenience: typically idempotency, audit, and permission handling that low-code tools do not express well. ##### How do you handle systems with no usable API? By being honest about the trade-off. Scheduled exports, file drops, and database-level integration are all legitimate when an API is absent, provided the same guarantees hold: idempotency, dry runs, audit. Screen-scraping a system of record is normally where we would advise against automating at all. Last reviewed 27 July 2026. ### Fractional AI platform architect & technical lead (https://www.1aym.com/capabilities/fractional-ai-platform-architect) Reference: C-07. Question this page answers: How do we get senior AI architecture ownership without hiring a permanent platform leader? Sold as: Fractional AI Platform Architect. Retained · typically 2–3 days a week · one statement of work. A permanent AI platform leader takes two quarters to find, sign and land, and a programme without one makes its architecture decisions by accident in the meantime. A fractional engagement puts that ownership in place now. We hold the target architecture, the vendor and model decisions, the technical review of work already in flight and the delivery governance, and we enable your own engineers to take all of it over. It is retained rather than fixed-scope, typically two to three days a week under one statement of work, and every decision is written down and dated so the reasoning outlives the engagement. #### The problem this solves A programme with a budget, a mandate and nobody who owns the architecture. The decisions still get made. They get made by whoever is in the room: the loudest vendor, the team with the nearest deadline, or the last proof of concept that happened to work. A permanent hire fixes it eventually. Two quarters of search, notice and ramp is the usual cost, and most of the decisions that shape the platform are taken before that person arrives. #### How the engagement runs A standing cadence, so the ownership is real rather than an advisory line on a slide. - **Architecture ownership**: We hold the target architecture and the trade-offs behind it, and we are answerable for them the way an employed lead would be. - **A weekly decision log**: Each decision, the options weighed and the reason for the one taken, written down and dated. It is the artefact that outlasts the retainer. - **Review of work in flight**: Technical review of what your teams and your suppliers are already building, so a problem is caught while it is still a design question. - **Vendor and model decisions**: Which model, which platform, and where the build-versus-buy line falls, with the production evidence behind each choice stated rather than asserted. - **Enablement**: The point of the engagement is that it ends. Your engineers take the architecture, the log and the operating model, and run them without us. #### What you are left with An architecture and a decision record your own team can defend to a board, to an auditor or to the next supplier, and engineers who helped write both. A written handover, so the end of the retainer is a date in the contract rather than a cliff. #### When it is the wrong choice When the problem is already scoped, a fixed-scope engagement is cheaper and faster. A feasibility sprint settles what to build; an architecture and production build ships it. Both are priced against a defined output, which a retainer is not. It is also the wrong shape when what is missing is capacity rather than ownership. That is a contract for embedded engineers, and it has its own page below. #### What you get - Target architecture, owned and kept current - A dated decision log covering vendor, model and build-versus-buy choices - Technical review of work already in flight - Delivery governance: gates, review points and release criteria - Enablement so your own engineers take the architecture over - A written handover at the end of the retainer #### Start here if - A funded AI programme with no single owner of the architecture - Vendor recommendations are the only technical opinion in the room - A permanent platform lead has been open for a quarter or more - Work is in flight and nobody senior is reviewing the design #### Questions ##### How much of the week does a retained architect hold? Typically two to three days a week, under one statement of work. The shape matters more than the total: ownership that only appears at an escalation is advice rather than ownership, so the days are a standing cadence. ##### Who makes the final call on a model or a vendor? You do. We hold the decision, set out the options and the reasoning, and put a date on it. What that buys is a record you can defend a year later, and an argument settled on evidence from production rather than on who presented last. ##### What stops this becoming a permanent dependency? The enablement sits inside the engagement rather than after it, and the handover is written as we go. Your engineers hold the architecture and the decision log alongside us, so the end of the retainer is a date rather than a risk to manage. Last reviewed 22 August 2026. ### Embedded Engineers on Contract (https://www.1aym.com/capabilities/embedded-engineers-on-contract) Reference: C-08. Question this page answers: How do we add senior engineers to a programme that is short of capacity rather than short of a plan? Sold as: Embedded Engineers on Contract. From three months · priced per day. We place a team, or a single senior engineer, inside your programme on a contract basis. Contracts run from three months and are priced per day. Every engineer meets the same standard as the rest of the practice: certified on the platforms we build on, or shipped inside a top-tier engineering organisation. You direct the work day to day, in your backlog and your stand-ups. 1AYM holds the standard and the cover behind it, so a stream never hangs off one person. The usual roles are an AI or platform architect, an AI engineer, and a data or platform engineer. #### The problem this solves A funded programme waiting on a hiring round. The plan is signed off and the budget is committed, and the start date now depends on how long it takes to find people. Or a delivery team short one senior role, with the date unmoved. Or a backlog that needs engineers rather than another roadmap. None of those is a defined output, so buying one as a fixed-scope engagement fits badly. #### How it runs One contract, a minimum of three months, on day rate commercials rather than against a deliverable. - **Who turns up**: A single senior engineer or a small team, inside your programme. The usual roles are an AI or platform architect, an AI engineer, and a data or platform engineer. - **The standard**: Every engineer meets the same bar as the rest of the practice: certified on the platforms we build on, or shipped inside a top-tier engineering organisation. - **Who directs the work**: You do, day to day, in your own backlog and your own stand-ups. We hold the standard and the cover, so a stream never hangs off one person. - **Weekly review**: One of our architects reviews the work each week, so an embedded engineer is never the only senior pair of eyes on a design. - **Swap or scale**: Changing a role, or adding one, happens under the same contract rather than through a new procurement round. #### What you are left with The work, in your repositories and your environments, built the way your own engineers build. Nothing depends on a system only we can reach. A clean handover at the end, or a conversion to a fixed-scope build once the work is defined enough to price against an output. #### When a fixed-scope engagement is the better buy When the outcome is defined, buy the outcome. A feasibility sprint settles what to build, and an architecture and production build ships it. Both are priced against a deliverable rather than a day, which is the cheaper trade whenever the deliverable can be written down. When what is missing is the technical ownership rather than the hands, a retained architect is the closer fit. #### What you get - A named senior engineer, or a team, inside your programme - One contract, from three months, priced per day - Weekly technical review by a 1AYM architect - Work delivered in your repositories, environments and process - Cover across the stream, so no role is a single point of failure - A clean handover, or conversion to a fixed-scope build #### Start here if - A funded programme is waiting on a hiring round - A delivery team is short one senior role and the date has not moved - The backlog needs engineers rather than another roadmap - Capacity is the constraint, and the outcome is not yet scoped enough to price #### Questions ##### How quickly can somebody start? Typically weeks rather than a hiring round. We match the roles to the programme first, because a fast start in the wrong discipline costs more than a slower one in the right discipline, and we say plainly when we do not have the right person for a role. ##### What happens if an engineer is not the right fit? We swap them, under the same contract. That is the practical difference between a contract with a practice and a direct hire: the cover sits with us, so a stream is never left waiting on one person's notice period or one person's holiday. ##### Can an embedded engagement become a scoped build? Often, and it is usually the better buy once it can be. Three months inside the programme is enough to know what the output is worth and what it takes, which is exactly what a fixed-scope statement of work needs before anyone can price it honestly. Last reviewed 22 August 2026. ## Engagement files (https://www.1aym.com/work) Written at a public-safe level: capability, constraints and numbers. Clients are not named; internal project names and proprietary business logic are held back. ### Self-serve finance answers, without breaking permissions (https://www.1aym.com/work/finance-data-access-connector) Reference: D-06. Client: Client withheld. Delivered by 1AYM. A question about client or financial performance used to take around two hours to come back from the finance team. It is now answered in the meeting where it is asked, and nobody gained access they did not already have: the connector carries each person's existing Looker entitlements, so people see exactly what they were already entitled to see. Connecting the data took an afternoon. The week that followed went on agreeing, with the finance director, the definitions that make an answer correct, and that week was the actual work. - ~2 hrs → in the meeting · Time to a governed finance answer - 1 afternoon · To connect the data - 1 week · To make the numbers trustworthy - Unchanged · Existing permission model #### The problem Financial information sat behind the finance team. Anyone elsewhere in the business who needed a figure, for a client conversation or a meeting or a decision, raised a request and waited, typically a couple of hours. The finance team spent a meaningful share of its week answering questions that were, in principle, already answered by data the company held. The obvious fix, letting people query the data directly, ran straight into the reason the gate existed: financial data is not uniformly shareable. Some of it is open to the business, some of it is privileged, and any solution that flattened that distinction was worse than the queue. #### What was built A connector between the company's AI workspace and its existing Looker setup, so questions asked in natural language resolved against the same governed model the business already used for reporting. - **Permission inheritance**: The connector carries the user's own Looker entitlements. Someone with privileged access keeps it; someone without sees only what was already open to them. No parallel permission model was created, because a second model is a second thing to get wrong. - **Grounded in the existing model**: Answers resolve against the company's canonical metrics rather than against raw tables, so the figures match the ones finance would have given. - **Semantic layer**: The definitions that make an answer correct (what counts as a client, a period, a booked figure), expressed once so every question inherits them. #### The hard part was not the integration Connecting the data took a single afternoon. That speed was itself the problem: it confirmed a belief held across the business that this kind of work is plug-and-play. It is not. A connector that returns a number is trivial. A connector that returns the *same* number the finance director would have given you is a week of work, and that week is the entire value. We spent it with the finance director, rapidly iterating the semantic layer and checking the agent's answers against his, until the two agreed on definitions the business actually uses. Had the project stopped after the afternoon, it would have shipped something that looked finished and quietly disagreed with finance, which is worse than the queue it replaced, because the queue was at least right. #### Outcome People pull the figures they need in the meeting where the question comes up, rather than preparing a request and waiting on someone else's queue. The finance team stopped being a lookup service for questions the data could answer on its own. The wider result was a corrected assumption. The engagement demonstrated to leadership that the gap between a working connection and a trustworthy one is where the actual engineering lives, which changed how subsequent AI work at the business was scoped. Stack: Looker, Semantic layer, AI workspace connector, Permission inheritance. Last reviewed 27 July 2026. ### An IELTS speaking assessment built as a measuring instrument (https://www.1aym.com/work/education-ai-enablement) Reference: D-02. Client: A government-backed EdTech in the Middle East. Delivered by 1AYM. We hold end-to-end technical ownership of a government-backed EdTech's live production estate, and its IELTS speaking assessment is the centre of it. Most AI speaking scorers are one model call wrapped in a product; this one is a measuring instrument. It listens to the whole interview, eleven to fourteen minutes of it, gathers acoustic and linguistic evidence independently, and applies the official band descriptors in ordinary code, so every band can be reproduced and inspected line by line. Where it is not confident it says so and routes the case to a human examiner. The pipeline is complete and independently verified, and accent fairness and agreement with human examiners are release gates whose validation is scheduled rather than results we are claiming. - 11–14 min · The whole interview is scored, not a sample - 5 layers · Of evidence, and the AI never does the arithmetic - ×3 · Examiner model runs, median taken, escalate on disagreement - ≥ 0.70 · Agreement with human examiners: a release gate, not a result - ~1,170 · Automated tests on the speaking pipeline #### Why this is the hard problem Two trained human examiners marking the same IELTS speaking interview agree at a correlation of about 0.90 [1]. That is the ceiling. An automated marker is measured against a standard that people do not hit perfectly themselves, so the useful question is not whether it is right every time. It is whether you can see how it reached a band, and catch it when it is wrong. The regulatory floor moved while this was being built. ETS's 2026 update to the TOEFL marks all eleven speaking items on the test as AI scored [2]. Ofqual fined Cambridge English £875,000 over automated-marking errors that ran undetected for more than two years [3]. Under the EU AI Act, a system that evaluates learning outcomes is high-risk by classification rather than by anyone's opinion of it [4]. So the design problem was never how to get a model to output a band. It was how to build something an examiner, a regulator and a candidate's appeal can all read. #### The five layers The system listens to the whole interview rather than sampling it, and each layer produces evidence the next one can check. The AI never does the arithmetic. - **1 · Capture**: Transcription with word-level timestamps, hesitations preserved rather than tidied away. A fluency judgement that cannot see where somebody paused is a guess. - **2 · Acoustic measurement**: Speech rate, pause length and position, where hesitation falls, pitch range: measured deterministically in signal-processing code. These are the feature families ETS SpeechRater has used for over a decade [5], and they are arithmetic rather than opinion, so the same audio gives the same numbers every time. - **3 · Pronunciation**: A dedicated phoneme-level model on the Goodness-of-Pronunciation method, trained against a corpus scored by five expert raters. General-purpose multimodal models were tested for this job and were not good enough at phoneme and stress judgement, at 0.21 against 0.61 to 0.74 for specialist models [6]. - **4 · The examiner model**: One pass over the whole interview with every measured value in front of it, applying the official band descriptors. It runs three times, the median is taken, and disagreement between the runs routes the case to a human without anyone asking. It sits behind a swappable interface, on OpenAI audio models with a Gemini fallback, so the instrument does not depend on one vendor staying still. - **5 · Scoring and guardrails**: The four criteria are averaged, rounded by a documented rule and capped where an answer is off topic, all of it in plain code. Integrity checks for memorised answers, read-aloud delivery and synthetic voice raise a review flag; none of them silently alters a score. #### Confidence routing, and the number the design is built around Cambridge publishes the figures that make the case for this shape. Its Linguaskill automarker awarded the same CEFR grade as the examiners on 56.8% of tests marking alone, and on 95.6% under the hybrid model, where a response the computer is not confident about goes to a human [7]. Escalation is what moves that number, not a better model. So the instrument is built around knowing when to stop. Roughly 15% of interviews route to a human reviewer, and a further 5% are audited at random whether the system was confident or not, because a gate tested only on the cases it flagged tells you nothing about the ones it let through. Every reviewed case comes back as labelled data for the next version. #### Fairness and accuracy are release gates These are gates, not results. Stating them the other way round is the failure this whole design exists to avoid. - **Accent fairness**: The pronunciation layer's output is evidence, never the score. Score differences by candidate first language are tested as a formal gate, within 0.10 standard deviations, and a build that fails it does not ship. - **Agreement with examiners**: Quadratic-weighted agreement of at least 0.70 with human examiners, the threshold Williamson, Xi and Breyer set out for automated scoring [8]. - **The validation is scheduled**: The corpus is 300 interviews, each marked by two certificated examiners with a third resolving disagreements. It is booked rather than finished, so this page publishes the gates and no number against them. An accuracy figure without its validation behind it is the first thing a regulator would take apart. #### Where the build stands The pipeline is complete and has been independently verified. Around 1,170 automated tests cover the speaking pipeline, among them a regression suite where disabling any one guardrail breaks specific frozen cases, so a guardrail cannot be quietly dropped without a test going red. The core mathematics was hand-verified against worked examples rather than only against itself. Seven regulator-facing governance documents are drafted, and the rollout is staged behind controls rather than switched on. #### How it started The brief was a frontend. Building it meant living inside the product, and the mock test made the ceiling clear: it could tell a learner whether an answer was right, but not why, and not in a language the learner was comfortable being taught in. We built an AI tutor that teaches English in the learner's own language, and the engagement widened from there into the question the CEO actually had, which was where AI belonged across the company and in what order. #### The estate underneath The engagement is now end-to-end technical ownership of the live production estate: the exam platform, the APIs and the data. On 21 August 2026 the production database held 3,141 live student records, 1,086 of them added in the previous ninety days, and 393 mock assessments completed since November 2025. Those are counts from the system, not a projection. The database was replatformed into Google Cloud's Doha region to meet Gulf data-residency requirements, and the automated test estate across the platform went from around 300 tests to more than 1,400. #### Questions ##### How accurate is the automated IELTS speaking score? No accuracy figure is published for this system. Accent fairness within 0.10 standard deviations by candidate first language, and quadratic-weighted agreement of at least 0.70 with human examiners, are release gates the build has to pass. The validation corpus behind them is 300 interviews, each marked by two certificated examiners with a third resolving disagreements, and it is booked rather than finished. ##### How much of the interview is scored? All of it. The whole interview, eleven to fourteen minutes, passes through five layers of evidence rather than a sample. The AI never does the arithmetic: the four criteria are averaged and rounded by a documented rule in plain code. ##### What happens when the system is not confident? Roughly 15% of interviews route to a human examiner, and a further 5% are audited at random whether the system was confident or not, because a gate tested only on the cases it flagged tells you nothing about the ones it let through. Every reviewed case comes back as labelled data for the next version. #### Sources [1] IELTS test statistics: inter-rater reliability for Speaking: https://ielts.org/researchers/our-research/test-statistics [2] ETS: TOEFL iBT 2026 update, test blueprint and specifications: https://www.ets.org/content/dam/ets-org/pdfs/toefl/toefl-ibt-test-specifications-2026.pdf [3] Tes: Ofqual fines Cambridge English £875,000 over automated marking errors: https://www.tes.com/magazine/analysis/general/ofqual-fines-cambridge-english-ps875000-over-automated-marking-errors [4] EU AI Act, Annex III: high-risk systems, education and training: https://artificialintelligenceact.eu/annex/3/ [5] ETS SpeechRater: automated scoring of spoken responses in the TOEFL iBT test: https://www.ets.org/speechrater.html [6] Exploring the potential of large multimodal models as effective alternatives for pronunciation assessment, arXiv:2503.11229: https://arxiv.org/abs/2503.11229 [7] Cambridge English: Linguaskill, building a validity argument for the Speaking test, June 2020: https://www.cambridgeenglish.org/fr/Images/589637-linguaskill-building-a-validity-argument-for-the-speaking-test.pdf [8] Williamson, Xi and Breyer, A framework for evaluation and use of automated scoring, 2012: https://doi.org/10.1111/j.1745-3992.2011.00223.x Stack: Speech assessment, Signal processing, OpenAI audio models, Gemini, Postgres 16, Google Cloud, Multilingual LLM tutoring, AI roadmap. Last reviewed 22 August 2026. ### AI, data and automation enablement across a global agency (https://www.1aym.com/work/ai-platform-enablement) Reference: D-01. Client: A PE-backed marketing agency. Delivered by 1AYM. Sessions on this agency's AI platform now cost 60% less on average than they did before we re-architected its skills estate. An earlier change, a per-user API call replaced with one ten-minute sync, carries a forecast saving of £2–4M a year, modelled on token consumption and reviewed by the client's finance team. We hold the lead architect role for the platform and work alongside the agency's own finance, systems, data warehouse and AI tooling teams, on an AI programme that is theirs rather than ours. - 60% · Lower cost per session after the re-architecture - £2–4M · Forecast annual saving, finance-reviewed - ~600 · Skills on the platform we contribute to - 80%+ · Skills that exceeded description limits before the rebuild - 4 · Finance systems integrated - Ongoing · Engagement status #### Built with their teams, not around them Worth stating plainly, because case studies routinely blur it: the AI programme at this agency is theirs. Their leadership set the direction, their teams author the skills, and adoption across the business is their achievement. 1AYM is engaged on the engineering underneath, working alongside their finance, systems, data warehouse and AI tooling teams. That is the point rather than a caveat. A supplier who builds in isolation leaves behind a system only they understand, and the client discovers the dependency the first time something breaks after the invoice is settled. Building alongside their people means upskilling them as the work goes. The finance director who shaped the semantic layer with us can defend and extend its definitions himself, and the team maintaining the provisioning system understands why it resolves identity the way it does rather than treating it as a box that must not be touched. The measure of this kind of engagement is not what runs while you are there. It is what the client can still maintain, change and build on once you have gone. #### How the scope grew The engagement started as integration work. It expanded because the problems worth solving kept turning out to sit between teams rather than inside one: a finance reconciliation issue that was really a data contract issue, an identity request that was really a provisioning architecture issue. The value of being forward-deployed is being able to follow a problem across those boundaries instead of handing it off at each one. Over time the role became a bridge between finance, systems, the data warehouse, and the internal AI tooling teams. #### What the work covered - **Finance systems enablement**: Data warehouse and finance systems work across Xero, NetSuite, Zoho and QuickBooks: financial detail, mapping, reconciliation, management-accounts and trial-balance checks, and year-end rollover workflows. - **Identity and provisioning**: Integration pipelines with deterministic identity resolution, dry-run modes, diff-based writes, auditability and safe rollback patterns. - **Internal AI skill platform**: Engineering contributions to an internal Claude Code skill platform that now carries around 600 production skills. The skills themselves are authored across the business by the people who do the work; our part is the plumbing and reliability beneath them. - **Verifier-gated automation**: Managed-agent workflows for finance-critical and high-stakes operational tasks, built so deterministic checks decide what proceeds. - **Grounded analytics**: AI-enabled analytics workflows, including MCP-style patterns connecting the data warehouse and BI layer so AI answers resolve against canonical business metrics. - **Platform reliability**: Production support across Cloud Run and BigQuery/Airflow, deployment fixes, and the ongoing reliability work that keeps daily runs predictable. #### Optimisation nobody asked for The skills library lives in Notion. As written, every user's client fetched from the Notion API directly, which works perfectly at ten users and becomes a problem at a thousand, because the request volume scales with people multiplied by how often they work, against an API with rate limits and an availability budget that was never sized for it. We replaced that with a single synchronisation running every ten minutes, so the platform reads from a local copy rather than every client hitting the source. Request volume stopped scaling with headcount, and the dependency on Notion being fast and reachable at the exact moment someone works went away with it. The forecast saving from that one change is £2–4 million over a year, alongside a system that is materially more reliable. A number that size deserves its provenance: we modelled it on token consumption, and it was reviewed and accepted by a member of the client's finance team rather than asserted by the person who made the change. The arithmetic is not complicated. Under the old pattern, consumption scaled with users multiplied by working sessions multiplied by fetches per session, and at roughly a thousand people that product grows fast, because every one of those fetches pulled content into a context window that someone was paying for. Under the new pattern the cost is fixed: one sync every ten minutes, whatever the headcount is doing. This was not in a brief. It is the kind of thing you only see from inside the system, and the kind of thing worth raising rather than waiting for it to become an incident. #### Phase two: re-architecting the skills estate Sessions on the platform now come back faster and, by our own measurement, cost 60% less on average than they did before the rebuild. Across the number of people using it and the sessions each of them runs in a day, that compounds into a substantial annual saving. 1AYM now holds the lead architect role for the platform, in addition to the sync described above. Phase two happened because phase one worked. Adoption spread, the estate of skills grew with it, and the platform slowed under its own success. The skills had been written as flat, monolithic files, and over 80% of them exceeded the description limits, so every session carried far more context than the task in front of it ever used. The rebuild followed the structure the Anthropic SDK already defines, rather than a shape of our own. - **Skills split to the SDK's structure**: Each skill is split the way the SDK expects, with reference files the model loads on demand instead of one file loaded in full at the start of every session. - **Context injected, not carried**: Custom MCP servers that inject the targeted context a task actually needs, in place of whole files travelling through the session whether or not anything reads them. - **End-to-end sync**: One pipeline from the authoring tool to the running platform, so what an author publishes is what a session gets. - **Simpler to use than it was**: We worked directly with the client's team on how people interact with the platform, so the rebuild left it easier to work with rather than only cheaper to run. #### Why finance work sets the standard Finance-critical automation is the most demanding environment to build in, because there is no acceptable error rate that gets waved through. A reconciliation that is nearly right is a reconciliation that is wrong, and it will be found by an auditor rather than a user. Everything built here inherits that constraint: idempotent operations, dry-run modes, diff-based writes, audit trails, and rollback paths as defaults rather than hardening added later. Stack: Xero, NetSuite, Zoho, QuickBooks, BigQuery, Airflow, Cloud Run, Notion, SCIM, Claude Code. Last reviewed 22 August 2026. ### Identity provisioning for 1,000+ users across 600+ groups (https://www.1aym.com/work/identity-provisioning-at-scale) Reference: D-04. Client: A PE-backed marketing agency. Delivered by 1AYM. When somebody leaves, their access to Notion and Claude ends with their HR record rather than when a colleague remembers. We built the provisioning system that does it from scratch, because the client's IT team was over capacity and off-the-shelf group tooling did not fit a model where one person belongs to many groups: it covers more than 1,000 users across over 600 groups and has run for seven months. A membership matrix that size was never going to be kept right by hand. - 1,000+ · Users provisioned - 600+ · Groups managed - 7 months · Running in production - HR system · Single source of truth #### The problem Two tools needed access managed across the whole organisation, against a membership model where one person belongs to many groups and group membership changes as people join, move and leave. Done manually, a matrix of that size is not merely tedious. It is unreliable in a specific and dangerous way. Manual provisioning fails safe on joiners, because someone complains when access is missing. It fails unsafe on leavers, because nobody complains about access that should have been removed. That asymmetry is how organisations accumulate active accounts belonging to people who left months ago. #### Why it was built rather than configured The internal IT team was over capacity and did not have the time to take it on, which is the ordinary reason this kind of work does not get done anywhere. Off-the-shelf group tooling was considered and rejected. The organisation's membership model (overlapping groups, HR as the source of truth, two target systems with different provisioning semantics) did not fit the shape those tools assume, and bending the model to fit the tool would have left the mismatch to be handled manually anyway. #### How it works - **HR as the source of truth**: Provisioning is driven from the HR system, so joining, moving and leaving are reflected without a separate administrative step that someone has to remember. - **Deterministic identity resolution**: Matching a person across systems is done by explicit rules rather than fuzzy heuristics, because a wrong match grants the wrong person access. - **Diff-based writes**: Each run computes the difference between intended and actual state and applies only that, so a run is idempotent and a retry cannot compound. - **Dry-run mode**: Changes can be inspected before they land, which is necessary when a single run can alter access for hundreds of people. - **Audit trail and rollback**: What changed, for whom, and why, with a defined path back from a bad run. #### Outcome The system has run for seven months. Access reflects the HR system rather than someone's memory, joiners are provisioned without a ticket, and leavers lose access because a record changed rather than because somebody noticed. This was built inside the wider enablement engagement at the same organisation, and shares its engineering defaults: idempotency, dry runs, auditability and rollback are the baseline rather than additions. Stack: SCIM, Notion, Claude, HR system integration, Deterministic identity resolution. Last reviewed 27 July 2026. ### A regulated product that was not going to ship, shipped. (https://www.1aym.com/work/clinical-trials-platform-rescue) Reference: D-05. Client: A clinical-research software start-up. Delivered by 1AYM. A clinical-research software start-up had a part-built platform, no engineering environment its own team could ship from, and a launch that was not going to happen on the path it was on. We worked with the founders: first a development environment, with repositories, environments, continuous integration, review gates and a release discipline, then completion of the platform, then delivery to launch alongside their team. The product is live and in market, and the environment it ships from belongs to the company rather than to us. - Founders · Engagement partner - Dev environment · Built for the team to ship - Completed · Platform modules brought to release - Live · Launched and in market #### Where it was The company had built part of a product and could not get it out. The platform existed in pieces, there was no engineering environment that could carry a change from somebody's branch to a release anyone would trust, and on the path it was on it was not going to launch. That is a common shape and it is rarely a shortage of code. What is missing is the ordinary machinery a regulated product needs before a release can be a routine event: somewhere for the work to live, environments to run it in, checks that run on every change, and a repeatable way to put a version in front of users. Without it, every release is an act of nerve, and a team that is nervous about releasing stops releasing. #### What we put in place In this order, because none of the later pieces is safe without the first one. - **A development environment**: Repositories, environments, continuous integration and review gates, so a change is written, reviewed, tested and promoted the same way every time. Built for the founders' own team to work in rather than for us to operate on their behalf. - **Release discipline**: How a version is cut, what has to pass before it moves, and what happens when something fails: defined, and in use. Discipline that holds only while the supplier is in the building is not discipline. - **Completion of the platform**: With a path to release in place, the remaining modules were built out and taken to a state that could ship rather than to a state that could be demonstrated. - **Delivery to launch**: We took the platform to launch with the founders rather than handing over a repository and a document. A first release is where everything the process missed turns up, so it is the release worth being in the room for. #### Why it matters in regulated research Clinical research software is not judged only on whether it works. It is judged on the record it leaves behind: who did what, under whose delegated authority, against which version of which document, and whether all of that can be produced on request. The platform holds those systems on one inspection-ready record rather than in separate products that have to be reconciled afterwards. Those systems are trial management (CTMS), the electronic trial master file (eTMF), the investigator site file and the pharmacy site file (eISF and ePSF), quality management (eQMS), corrective and preventive actions (CAPA), and digital delegation of authority. It is built for NHS and hospital R&D, academic research, contract research organisations and mid-market sponsors, and its documentation is aligned to MHRA, FDA and EU Annex 11 expectations. That is the product's own public description of what it is for, and this page repeats it at that level. #### What the founders are left with The product is live and in market. A launch that was not going to happen on the previous path happened on this one, and it happened in a way the company can repeat. What stays behind counts as much as what shipped. The development environment is theirs and their own team runs it, the release discipline is theirs and in use, and the next release goes out through the same path that shipped this one. That is the measure we apply to every engagement on these pages: not what runs while we are there, but what the client can maintain, change and build on once we have gone. Stack: Development environment, Continuous integration, Review gates, Release management, CTMS, eTMF, eQMS and CAPA, Regulated documentation. Last reviewed 22 August 2026. ## Reference ### What is an AI harness? (https://www.1aym.com/what-is-an-ai-harness) Definition: A structured platform layer that connects models, tools, workflows, business context, and permissions into one coherent, governed operating layer, giving an organisation reusable internal AI capability instead of one-off agents. An AI harness is a structured platform layer that connects models, tools, workflows, business context, and permissions into one coherent, governed operating layer. Instead of each team building one-off agents and fragmented experiments, a harness gives an organisation reusable internal AI capability: shared skills, routing, governance, and feedback loops that every business vertical can build on. The harness is the thing that persists; individual agents and use cases are what you build on top of it. #### The problem a harness solves Most organisations do not have an AI problem. They have an AI sprawl problem. A team in finance wires a model into a spreadsheet, someone in operations builds a chatbot, a data team runs retrieval experiments, and none of it shares a permission model, an evaluation method, or a definition of what the business actually means by a customer. Each of those efforts works in isolation and none of them compounds. The second use case costs as much to build as the first. Nothing learned in one carries into the next. When something goes wrong, no one can say which component produced the output or whether it was allowed to. #### What sits inside a harness A harness is not one piece of software. It is the set of shared layers that every AI use case in the organisation draws on, built once and governed centrally. - **Model and tool routing**: Which model handles which step, which tools it may call, and how requests fall back when a model is slow, unavailable, or the wrong fit for the task. - **Business context**: The organisation's own definitions, data contracts, and domain vocabulary, expressed once so every agent inherits the same meaning of a customer, an invoice, or an active account. - **Permissions and identity**: What each agent may read and write, acting on behalf of which user, under which approvals, enforced by the platform rather than reimplemented per use case. - **Evaluation and feedback**: Shared harnesses for measuring output quality, regression suites that run when prompts or models change, and a route for human corrections to feed back into the system. - **Observability and audit**: A record of what ran, on whose behalf, against which inputs, producing which output: the thing that makes an AI programme answerable to a risk review. #### What an AI harness is not It is not a chatbot, and it is not a model subscription. Buying seats for a frontier model gives individuals a better tool; it does not give the organisation a capability. The difference shows up the moment you need a workflow to run unattended, touch a system of record, or survive an audit. It is also not a framework you install. The routing, the context layer, and the permission model all encode decisions specific to your business. Frameworks help you build a harness; they are not one. #### When an organisation needs one The signal is usually the third use case. The first two can be built directly and shipped. By the third, the cost of not having shared context, evaluation, and permissions becomes visible: the same integration work is being repeated, and no one can answer governance questions consistently across the three. The other common trigger is a use case that touches money, identity, or a regulated process, where 'mostly correct' stops being acceptable and the verification layer has to be engineered rather than assumed. The cost of leaving it is measurable rather than theoretical: on one platform we hold the lead architect role for, over 80% of skills had grown past the description limits before the rebuild, and re-architecting the estate cut the average cost per session by 60%. That is harness debt, priced. #### Questions ##### Is an AI harness the same as an agent framework? No. An agent framework is a library for constructing agents: orchestration primitives, tool-calling conventions, memory abstractions. An AI harness is the organisation-specific platform layer built above that: your business context, your permission model, your evaluation suites, your routing decisions. A framework is a component you might use inside a harness; it is not a substitute for one. ##### How long does it take to build an AI harness? It is built incrementally alongside real use cases, not as a platform project that precedes them. The usual route is to deliver one production use case, then extract the shared layers (context, permissions, evaluation) as the second and third are built. Building the platform first, in the absence of a working use case, tends to produce abstractions the business never needed. ##### Does an AI harness lock us into one model provider? No, the opposite, when built correctly. Routing is a layer the harness owns, so model choice becomes configuration rather than architecture. Last reviewed 22 August 2026. ### Agent proposes, verifier gates (https://www.1aym.com/agent-proposes-verifier-gates) Definition: An operating pattern for dependable AI automation in which the AI performs flexible reasoning while deterministic validators, schemas, tests, and human review gates decide whether the output is safe to proceed. “Agent proposes, verifier gates” is an operating pattern for dependable AI automation: the AI performs the flexible reasoning (research, drafting, transformation, generation) while deterministic validators, rules, schemas, tests, and human review gates decide whether the output is safe to proceed. The model is never the last thing to touch a decision. It is what makes automation dependable in finance, data migration, and compliance workflows where “mostly correct” is not good enough. #### Why “mostly correct” fails A model that is right 95% of the time sounds excellent until you run it a thousand times a day against a ledger. Then it is fifty wrong entries a day, arriving with the same confident tone as the correct ones, in a system where the cost of a wrong entry is not one-twentieth of the value of a right one. Accuracy is the wrong frame for automation. The question is not how often the model is right. It is what happens on the occasions it is wrong, and whether the system notices before the consequence lands. #### How the pattern works The work is split by what each component is actually good at. Models are good at ambiguity, language, and synthesis. Deterministic code is good at being certain. The pattern uses each for what it is reliable at, and never asks the model to certify itself. - **The agent proposes**: It reads the ticket, drafts the entry, maps the fields, extracts the values, or writes the summary: the part that requires judgment across messy inputs. - **The verifier gates**: Schema validation, business rules, reconciliation against a source of truth, type and range checks, and tests. None of it involves a model. All of it can be reasoned about and unit tested. - **Pass proceeds, fail holds**: Passing output continues automatically. Failing output routes to human review with the reason attached, rather than being silently retried or silently shipped. - **The gate is the audit trail**: Because every output passes an explicit check, the record of what was checked and why it passed is a by-product of running the system, not extra compliance work bolted on afterwards. #### What the verifier actually checks The strongest verifiers are boring. They assert that a total reconciles, that an identifier exists in the system of record, that a date falls inside the period, that a required field is populated and typed correctly, that a proposed write is idempotent. Where a deterministic check is genuinely impossible, the gate becomes a human one, but a narrow one, presented with the specific claim to confirm rather than the whole task to re-do. The aim is to spend human attention only where it is irreplaceable. #### Where the pattern applies It earns its keep anywhere the cost of a silent error exceeds the cost of a held item: finance-critical automation, data migration validation, knowledge extraction, semantic-layer query validation, report generation, operational triage, and internal tooling that writes to systems of record. It is unnecessary where output is inherently reviewed by the person who asked for it: a drafting assistant a human reads before sending does not need a gate, because the human is one. #### Questions ##### Does this slow the system down? Not meaningfully. A deterministic check costs microseconds against a model call's seconds. What changes is that some items stop instead of completing, which is the job. ##### Can the verifier be another model? Sometimes, but it is a weaker guarantee and should never be the only gate on a consequential action. A model checking a model shares failure modes with it, and both can be confidently wrong about the same thing. Model-based evaluation is useful for scoring quality at scale; deterministic validation is what you use to decide whether something is safe to write. ##### How is this different from a human-in-the-loop workflow? Human-in-the-loop usually means a person reviews everything, which does not scale and quietly degrades into rubber-stamping. This pattern routes only what fails an explicit check to a person, with the reason attached, so human attention is spent on genuine exceptions rather than spread thin across a queue that is mostly correct. Last reviewed 27 July 2026. ## Credential register (https://www.1aym.com/claude-certified-architect) The credential register: both Claude certifications held by 1AYM’s founder and Principal Architect, rendered verbatim from Credly’s public record. ## Every route - https://www.1aym.com/ · 1AYM homepage - https://www.1aym.com/about · About 1AYM - https://www.1aym.com/capabilities · Services - https://www.1aym.com/capabilities/executive-ai-discovery · Executive AI discovery & product translation - https://www.1aym.com/capabilities/ai-harness-platform-engineering · AI harness & proprietary platform engineering - https://www.1aym.com/capabilities/production-ai-systems · Production AI systems - https://www.1aym.com/capabilities/agentic-workflow-design · Agentic workflow design - https://www.1aym.com/capabilities/data-platform-ai-enablement · Data platform & AI enablement - https://www.1aym.com/capabilities/enterprise-integrations-automation · Enterprise integrations & automation - https://www.1aym.com/capabilities/fractional-ai-platform-architect · Fractional AI platform architect & technical lead - https://www.1aym.com/capabilities/embedded-engineers-on-contract · Embedded Engineers on Contract - https://www.1aym.com/work · Engagement files - https://www.1aym.com/work/finance-data-access-connector · Self-serve finance answers, without breaking permissions - https://www.1aym.com/work/education-ai-enablement · An IELTS speaking assessment built as a measuring instrument - https://www.1aym.com/work/ai-platform-enablement · AI, data and automation enablement across a global agency - https://www.1aym.com/work/identity-provisioning-at-scale · Identity provisioning for 1,000+ users across 600+ groups - https://www.1aym.com/work/clinical-trials-platform-rescue · A regulated product that was not going to ship, shipped. - https://www.1aym.com/claude-certified-architect · Claude Certified Architect - https://www.1aym.com/what-is-an-ai-harness · What is an AI harness? - https://www.1aym.com/agent-proposes-verifier-gates · Agent proposes, verifier gates - https://www.1aym.com/privacy · Privacy notice - https://www.1aym.com/llms.txt · llms.txt - https://www.1aym.com/llms-full.txt · llms-full.txt ## Contact - Email: tayyeb@1aym.com - Book a 30-minute call: https://cal.com/tayyeb-mahmud-jxedia - Sitemap: https://www.1aym.com/sitemap.xml - Short form of this file: https://www.1aym.com/llms.txt Clients are referred to by descriptor in this file and on every route, with two exceptions: the engagement file whose published source names its client, and the homepage logo bar, which shows two client marks. Descriptors are used everywhere else.