Guide

How much does AI implementation cost, and what drives the number?

AI implementation cost is everything it takes to get an AI use case from an idea to a system people rely on, and to keep it running. A budget has six lines: discovery, build, integration with the systems the work touches, evaluation and controls, change management, and run costs. The total moves with five things: how many systems the work connects to, how ready the data is, how much checking the output needs, how many people have to change how they work, and how often it runs. For one workflow at moderate volume, people’s time is usually the largest line and model usage one of the smallest. For a public US reference point, GSA’s federal schedule listed 204 AI and machine learning labor categories on 29 September 2026, with a median hourly ceiling rate of $190.74.

AI implementation cost. The total cost of taking an AI use case from idea into daily use and keeping it running: discovery, build, integration, evaluation and controls, change management and run costs.

1AYM is an OpenAI Select Partner. Its founder holds personal Claude certifications. 1AYM is not an Anthropic partner.

Checked . Federal rates, US wage and compensation figures, the accounting standard and vendor prices were read on the publishers’ own pages. Rates and prices change, so confirm them on the sources below before you rely on them. None of the figures on this page is a 1AYM price.

What an AI implementation budget pays for

Search for what AI development costs and you get wide ranges, mostly published by firms that sell AI development. I would rather give you the lines the number is made of and the public figures you can price them with, so the estimate is yours and you can defend it to a finance committee. Every external figure on this page comes from a primary source listed at the end, read on the date above, and none of them is 1AYM’s price.

The build is the line everyone asks about, but a budget is only as sound as its least-scoped line: an integration nobody scoped, a data clean-up nobody expected, or a system that works but that nobody changed their day to use.

The six lines of an AI implementation budget, what each one buys and what moves its cost
DimensionWhat it buysWhat moves the cost
DiscoveryChoosing the use case, measuring the work as it runs today, checking the data and systems, and deciding whether to build or buyHow many candidate use cases are on the table, and whether anyone can say what the work costs today
BuildThe workflow, prompts, retrieval, interfaces and the code around the modelHow many steps the model decides for itself, and how exact the output has to be
IntegrationConnections to the systems the work reads from and writes to, with identity and permissionsThe number of systems, whether each has a usable API, and whether the AI step writes or only reads
Evaluation and controlsA test set of real cases, checks on every consequential action, logging and approval routesThe cost of a wrong answer: a draft a person reads needs little, a write to money or customer records needs all of it
Change managementTraining, rewritten procedures, a named owner and the time people spend learning the new way of workingHow many people change how they work, and whether their managers measure the new way
Run costsModel usage, hosting, monitoring, seats, retesting when a vendor changes a model, and the owner’s timeVolume, how much text each run carries, and how often the vendor or the business changes something

The five things that move the number

Each driver below can be estimated before a build starts. If a supplier cannot tell you which of them sets their quote, the quote is a guess.

How many systems it touches
Every system adds connection work, permissions and testing. Reading from one system is the cheap case. Writing to three, one of them without a usable API, is a different project. Count the systems and mark each one read or write before you ask for a price.
How ready the data is
A model is only as right as the records it reads. If teams define a customer, an order or a margin differently, agreeing the definitions is part of the budget, and it can take longer than connecting the data. On one of our published engagements, file D-06, connecting finance data took an afternoon and agreeing the definitions with the finance director took the week that followed.
What a wrong answer costs
This sets how much checking you pay for. NIST’s AI Risk Management Framework, voluntary and released in January 2023, makes measurement one of its four functions alongside govern, map and manage. In budget terms that is a test set of real cases run on every change, and deterministic checks that decide what an AI step may change.
How many people change how they work
A system nobody uses has cost the full build and returned nothing. Training, rewritten procedures and a manager who measures the new way are budget lines, and they grow with the number of people whose day changes.
How often it runs
Model usage scales with runs a day and the text each run carries. Work that can wait costs less: OpenAI and Anthropic each publish a 50% discount on their batch APIs for requests that do not need an answer straight away.

What the US market charges for AI engineering time, from published rates

Firms that sell AI work rarely publish their prices. The public exception is the federal government’s Multiple Award Schedule, where GSA publishes the hourly ceiling rates it has awarded to each vendor’s labor categories, in its CALC+ tool. GSA’s user guide describes them as not-to-exceed prices it has determined fair and reasonable, and says they should not be viewed as exact estimates for a particular location, order size or complexity. An order cannot exceed them, and commercial prices can differ, so read them as reference points rather than quotes.

On 29 September 2026 we queried every labor category CALC+ returns for “artificial intelligence” and “machine learning”: 204 categories from 68 vendors, at their current-year hourly ceilings. The spread matters more than any one figure.

Hourly ceiling rates for AI and machine learning labor categories on GSA’s Multiple Award Schedule, read from CALC+ on 29 September 2026
DimensionCategoriesMedian hourly ceilingMiddle half of rates
All AI and machine learning categories204$190.74$146.55 to $244.95
Requiring three years’ experience or fewer65$148.72$115.29 to $198.76
Requiring ten years’ experience or more42$245.60$191.34 to $288.19
Small business vendors133$178.00$137.11 to $234.28
Other than small business vendors71$223.09$168.85 to $259.94

The full range ran from $55.60 to $485.47 an hour. At the median, a 40-hour person-week is $7,629.60 at ceiling.

Compare that with employing the skills yourself. The Bureau of Labor Statistics puts the May 2025 median pay for its combined group of software developers, quality assurance analysts and testers at $134,040 a year, or $64.44 an hour, and for data scientists at $120,230, or $57.80 an hour. Wages and salaries were 70.0% of private employers’ compensation costs in June 2026. Applying that all-industry share to $64.44, a rough loaded cost for that median is about $92 an hour, before recruiting, management, equipment and the months a hire takes to arrive.

Neither number is the answer on its own. An in-house team costs less per hour and more per idle month, because you pay for it whether or not the work is ready. An outside firm costs more per hour and, on a fixed scope, carries the risk of an overrun. What decides it is how long you need the skills and whether the work is defined well enough to price as an output.

A software company adding AI to the product it sells makes a narrower version of this choice, and our guide to AI features in a software product compares partnering, hiring and doing both.

Run costs: small per call, decided by design at scale

Model usage sounds alarming in the abstract and is often modest for a single workflow. Anthropic’s own pricing page works an example of 10,000 support conversations of about 3,700 tokens each on Claude Haiku 4.5, at $1 per million input tokens and $5 per million output tokens, for roughly $37 in total. Set beside a loaded engineering hour of about $92, that is small, so for one workflow at moderate volume the model bill is unlikely to be the line that decides the budget.

At enterprise scale the arithmetic changes, and the design decides it. At a large international marketing agency, one change, replacing a per-user API call with a single ten-minute sync, carries a forecast saving of £2–4M a year, modelled on token consumption and reviewed by the client’s finance team. The detail is in engagement file D-01. The saving came from the design, which is where run cost at scale is decided.

The run costs that surprise people sit around the model: monitoring, retesting when a vendor retires or changes a model, seats for the people using AI tools, and a named owner’s time every week. Budget them per year, and put a ceiling per run and per month into the design so a runaway loop cannot turn into an invoice.

Fixed scope, retainer or time and materials: who carries the overrun

The same work can be bought three ways, and the shape changes the risk more than the headline figure does. Buy the shape that matches how well the work is defined today.

Three ways to buy AI implementation work: what each one prices, who carries an overrun, and when each fits
DimensionWhat you pay forWho carries an overrunFits when
Fixed-scope statement of workA defined output, such as a scored shortlist of use cases or one workflow in productionThe supplier, within the written scopeThe output can be written down before the work starts
RetainerA standing share of a senior team’s week, over monthsShared: you set the priorities, the team owns the deliveryA programme is funded but nobody owns the architecture or the build
Time and materialsHours worked, at agreed hourly ratesYouThe work cannot be defined yet and you will direct it closely

Whichever shape you choose, ask three questions before signing: what exactly each fee pays for, what the system costs to run for a year after launch, and what you own if you stop paying. A small build fee followed by a monthly charge for workflows held in the supplier’s account can cost more over two years than a fixed-scope build you own outright.

Capitalise or expense: the accounting change US finance teams should plan for

For a US company, how the build is accounted for can matter as much as what it costs. In September 2025 the Financial Accounting Standards Board issued ASU 2025-06, which amends the guidance on internal-use software in Subtopic 350-40 and removes its old project stages. Under the new guidance, capitalisation starts when management has authorised and committed to funding the project and it is probable that the project will be completed and the software used to perform the function intended.

That threshold is not met while there is significant development uncertainty, which the standard ties to two questions: whether novel or unproven functions have been resolved through coding and testing, and whether the company has settled what it needs the software to do. An AI use case that is still an experiment may not clear either. My reading is that discovery gains a second job here, because the decisions it records are the kind of evidence an auditor looks for when that uncertainty is resolved.

FASB expects capitalisation to change little for most types of software, and says it could decrease for software developed to be provided through a cloud computing arrangement. The amendments are effective for all entities for annual reporting periods beginning after December 15, 2027, and interim periods within them, and early adoption is permitted as of the beginning of an annual reporting period. Whether an AI system counts as internal-use software, and how your costs are treated, is for your auditor. This is not legal, tax or accounting advice.

A budgeting worksheet to copy

Fill this in before you ask anyone for a quote. A line you cannot fill is discovery work you have not done yet. The reference rates at the bottom are the public figures above, with their date; replace them with your own quotes and payroll data as they arrive.

AI implementation budget worksheet
AI IMPLEMENTATION BUDGET WORKSHEET

Use case: ______________________
Business owner: ______________________
What the work costs today (hours a week, or cost a year): ______
The measure that must move, and the date you will read it: ______

ONE-OFF COSTS
1. Discovery: people ___ x weeks ___ x weekly rate $___ = $___
2. Build: people ___ x weeks ___ x weekly rate $___ = $___
3. Integration: systems ___ (mark each READ or WRITE) x weeks per system ___ x weekly rate $___ = $___
4. Evaluation and controls: real test cases to collect ___; checks to write ___; weeks ___ x weekly rate $___ = $___
5. Change management: people affected ___ x training hours each ___ x their loaded hourly cost $___ = $___
6. Contingency for what discovery could not settle: $___

YEARLY RUN COSTS
7. Model usage: runs a month ___ x cost per run $___ x 12 = $___
8. Hosting, monitoring and logging: $___ a month x 12 = $___
9. Seats for AI tools: users ___ x price per seat a month $___ x 12 = $___
10. Owner and support time: hours a month ___ x loaded hourly cost $___ x 12 = $___
11. Retesting when a vendor changes a model: times a year ___ x cost per retest $___ = $___

YEAR ONE TOTAL = lines 1 to 11
YEAR TWO RUN RATE = lines 7 to 11

REFERENCE RATES, 29 SEPTEMBER 2026 (replace with your own)
Outside firm, weekly: GSA CALC+ median AI and machine learning hourly ceiling $190.74 x 40 hours = $7,629.60 a person-week
In house, hourly: BLS median pay, software developers, QA analysts and testers, $64.44 / 0.70 wage share of compensation = about $92 loaded

If lines 3, 4 and 6 are the hardest to fill, the project is not ready to be priced as a build, and the honest next step is a short, fixed-scope discovery that fills them.

Where 1AYM fits

If the discovery lines of that worksheet are the ones you cannot fill, because nobody has yet measured the work, checked the data or settled build against buy, that is what our AI Opportunity & Feasibility Sprint is for: a fixed-scope engagement of typically two to four weeks that scores candidate use cases on value, feasibility and risk, and ends in a plan and an investment case your finance team can fund. After it, the build is a fixed-scope statement of work that can start within a day of the scope being signed, or, where a funded programme needs its architecture owned for months, a retained implementation team for typically two to three days a week.

We take small fixed-scope statements of work as well as larger builds, and US clients can contract through 1AYM’s US entity. We do not publish prices, because a price without a scope is the problem this page is about. A 30-minute call is enough to tell you which line of your worksheet we would start on.

For engineers: estimating run cost and build effort from the design

The run-cost lines in engineering terms, with the vendor mechanics that move them. Vendor prices were read on the vendors’ own pages on the date above and change often.

Cost per run
Sum over every model call in a run: input tokens times the input price, plus output tokens times the output price, plus the tool-definition tokens on each call that declares tools. Anthropic lists its tool-use system-prompt overhead per model. Measure your own prompts with a token-counting call rather than estimating from word counts.
Tokenizer drift
Token counts do not carry across model generations. Anthropic states that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, so a price comparison across generations needs a re-count on your own prompts.
Prompt caching
Put stable context, such as instructions, schemas and reference documents, first so it can be cached. On Anthropic’s API a cache read costs 0.1 times the base input price on most models and a five-minute cache write 1.25 times, so caching pays back after one read.
Batch work
Anything that can wait belongs in a batch. OpenAI’s Batch API carries a 50% discount against its synchronous APIs, with each batch completing within 24 hours. Anthropic’s Batch API is 50% off both input and output tokens.
US-only processing
If data must be processed in the US, price it in. Anthropic applies a 1.1 times multiplier to all token pricing for US-only inference on Claude 4.6 and later models. Check each vendor’s and cloud platform’s regional pricing for the model you choose.
Stop conditions
Cap model calls, tool calls and tokens per run, fail closed on a breach, and alert on spend per workflow per day. A loop without a ceiling is the one run cost with no upper bound.
Evaluation cost
A regression set of real cases runs on every prompt, model or tool change. Budget it as cases times runs a month times cost per case, plus the engineering time to keep expected answers current when the business changes its definitions.
Sizing integration
Estimate per system and per direction. A read through a documented API, a write with idempotency and a dry-run mode, and a system with no API at all are three different estimates. The third kind, found after the quote, is a common source of overruns.

Sources

  1. [1]GSA, CALC+ Quick Rate: hourly labor ceiling rates on the Multiple Award Schedule (data read 29 September 2026)
  2. [2]GSA, User Guide: CALC+ Quick Rate hourly labor ceiling rates, version 2.0 (24 June 2026)
  3. [3]U.S. Bureau of Labor Statistics, Occupational Outlook Handbook: Software developers (May 2025 wages; page modified 27 August 2026)
  4. [4]U.S. Bureau of Labor Statistics, Occupational Outlook Handbook: Data scientists (May 2025 wages; page modified 27 August 2026)
  5. [5]U.S. Bureau of Labor Statistics, Employer Costs for Employee Compensation, June 2026 (released 9 September 2026)
  6. [6]NIST, AI Risk Management Framework (AI RMF 1.0, released 26 January 2023)
  7. [7]FASB, Accounting Standards Update 2025-06, Intangibles, Goodwill and Other, Internal-Use Software (Subtopic 350-40) (September 2025)
  8. [8]Anthropic, Claude API pricing (read 29 September 2026)
  9. [9]OpenAI, Batch API guide (read 29 September 2026)

Frequently asked questions

How much does AI development cost?

There is no honest single figure, because the cost is the sum of six lines: discovery, build, integration, evaluation and controls, change management and run costs. Each moves with the systems touched, the state of the data, the cost of a wrong answer, the people affected and the volume. For a public US reference point, GSA’s federal schedule listed 204 AI and machine learning labor categories on 29 September 2026 with a median hourly ceiling of $190.74, or $7,629.60 for a 40-hour person-week. Multiply that by the weeks each line needs to reach a first estimate.

How much does AI consulting cost in the US?

Few firms publish prices. The public reference is GSA’s CALC+ tool, which lists the hourly ceiling rates awarded on the federal Multiple Award Schedule. For AI and machine learning labor categories on 29 September 2026 the median was $190.74 an hour, with the middle half between $146.55 and $244.95. These are ceilings for federal orders, not commercial quotes. A fixed-scope price for a defined output tells you more than any hourly rate, because it shows who carries an overrun.

What does an AI system cost to run each year?

Add model usage, hosting and monitoring, seats for the people using AI tools, the owner’s time, and retesting whenever a vendor changes a model. For one workflow at moderate volume the model bill is often small: Anthropic’s own worked example prices 10,000 support conversations on Claude Haiku 4.5 at roughly $37. At enterprise volume the design decides the bill, so set a cost ceiling per run and per month before launch.

Is it cheaper to build AI in-house or to hire a firm?

Per hour, in-house is cheaper: BLS put the May 2025 median pay for US software developers, quality assurance analysts and testers at $64.44 an hour, about $92 loaded with benefits, against a median federal ceiling of $190.74 for AI and machine learning categories. Over a year it depends on duration and definition. A permanent team costs money whether or not the work is ready, and takes months to hire. An outside firm on a fixed scope carries the risk of an overrun. Short, definable work favours a firm; a capability you need for years favours your own team, often built with outside help at the start.

Can AI development costs be capitalised under US GAAP?

Possibly, where the system is internal-use software. FASB’s ASU 2025-06, issued in September 2025, starts capitalisation once management has authorised and committed to funding the project and completion is probable, and not while significant development uncertainty remains. It is effective for annual reporting periods beginning after December 15, 2027, with early adoption permitted. Treatment depends on your facts, so ask your auditor. This is not legal, tax or accounting advice.

Why do AI projects go over budget?

An overrun comes from work that was not in the estimate: integrations discovered after the quote, data definitions that teams disagree on, checks added late because nobody priced the cost of a wrong answer, and adoption work left out of the budget. Each of those can be estimated before a build starts, which is what a scoped discovery phase is for.

Further

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM