Guide

AI spend governance: how to budget for seats, tokens and metered agent runs

AI spend governance is the budgets, limits and reports that keep AI costs predictable and tied to an owner. It starts with how the bill is built, and on the vendors' own pricing pages, checked on 29 September 2026, there are four meters. Seats are a fixed monthly price per person: US$20 a seat billed annually for a standard Claude Team seat or for ChatGPT Business. Allowances and credits set how much use a seat includes and what use beyond it costs. Tokens, the units of text a model reads and writes, are priced per million, from US$0.10 to US$10 per million input tokens across OpenAI's GPT-6 models. Agent runs can carry a meter of their own on top of tokens, such as US$0.08 per session-hour for Anthropic's Claude Managed Agents. Seats are easy to budget; the other three move with model choice, context and volume. Each needs a named owner, a limit that actually stops spending, alerts before it, and a monthly review against a baseline from your own pilot.

AI spend governance. The budgets, spending limits, allocation rules and reviews an organisation uses to forecast, cap and attribute what it pays for AI seats, usage credits, API tokens and metered agent runs.

1AYM is an OpenAI Select Partner. Its founder holds personal Claude certifications. 1AYM is not an Anthropic partner.

Checked . Prices are list prices in US dollars from OpenAI's and Anthropic's own pages, before any negotiated discount. Anthropic's plan page says its prices exclude applicable tax; the OpenAI pages we read do not say either way. Vendors change prices, plans, models and limits without notice, so check the sources below on the day you budget.

Four meters make up an AI bill

Before anyone can set a budget, finance needs to know which meter each line of the bill runs on. There are four, and one vendor can charge on several of them at once. Claude Enterprise, for example, is priced as a seat plus usage at API rates [2], so a company that thinks it has bought seats has also bought a metered service.

The four ways AI is billed, with examples in US dollars from each vendor's own pages, read 29 September 2026. List prices before negotiated discounts [1, 2, 3, 4, 5, 8].
DimensionHow it is billedExamples from the vendors' pagesWhat makes it move
SeatsA fixed price per person per month, whatever they useClaude Team: US$20 a standard seat or US$100 a premium seat a month, billed annually (US$25 and US$125 billed monthly) [2]. ChatGPT Business: US$20 per user a month billed annually, or US$25 monthly, for two or more users [1]. Claude Enterprise: US$20 a seat a month billed annually, plus usage at API rates [2]Headcount and seat tier
Allowances and creditsEach seat includes a usage allowance; work beyond it is paid for in credits at your plan's or agreement's rateClaude Team and Enterprise allowances reset on a rolling five-hour window and a weekly window, shared across chat, Claude Code and Cowork [8]. Codex and ChatGPT Work draw on one allowance, and credits past it are priced per million tokens [1, 5]A few heavy users, and which features they use
TokensA price per million tokens read (input) and written (output), with a lower rate for input the model has cachedGPT-6 Sol and Claude Sonnet 5.5: US$2 input and US$10 output per million tokens. GPT-6 Luna: US$0.10 and US$0.50. GPT-6 Astra: US$10 and US$50. Claude Opus 5.5: US$4 and US$20 [3, 4]Model choice, the context each call carries, output length and volume
Metered agent runs and toolsRuntime or per-call charges on top of the tokens an agent usesClaude Managed Agents: US$0.08 per session-hour of running time, plus tokens [4]. OpenAI's hosted shell and Code Interpreter containers: from US$0.03 per 20-minute session at 1 GB [3]. Web search: US$10 per 1,000 searches on either vendor's API, plus the tokens the results add [3, 4]How long agents run, how many tools they call and how many runs a day

Seats are the only meter a finance team can forecast from a headcount plan. The other three are consumption, and consumption follows behaviour: which model a team picked as its default, how long a session runs before anyone clears it, and whether a scheduled agent is still running for a project that ended last quarter.

Why the token line is the one that surprises people

A token price looks small, which is exactly why it misleads. The bill is that price multiplied by a quantity nobody decided, and several settings change the price itself. These are the ones I would check first, all read from OpenAI's and Anthropic's pricing pages [3, 4].

Model choice
Across OpenAI's GPT-6 models the output price runs from US$0.50 to US$50 per million tokens, a factor of 100 (our arithmetic). Anthropic's own advice is Haiku for simple tasks, Sonnet for most production work and Opus for the hardest reasoning [4]. Whoever sets the default model is setting the budget, usually without being told.
Output against input
Output costs five times input on GPT-6 Sol, Claude Sonnet 5.5 and Claude Opus 5.5 [3, 4]. Reasoning counts as output: Anthropic bills extended thinking as output tokens, and the default thinking budget can run to tens of thousands of tokens a request on some models [8].
Context carried on every call
Every call pays for everything the model is sent: conversation history, documents and tool definitions. Cached input is billed at a tenth of the input price on GPT-6 Sol and Claude Sonnet 5.5, so a stable prompt prefix that hits the cache is the cheapest saving on this list [3, 4].
Speed, batch and location
Both vendors' Batch APIs halve token prices for work that can wait. Fast mode costs twice the standard rate on GPT-6 Sol and Claude Opus 5.5. Keeping inference in the US costs 1.1 times the standard rate on Claude 4.6 and later models, and OpenAI adds 10% for regional processing on eligible models released on or after 5 March 2026 [3, 4].
Tokenizers
A price per million tokens only compares like with like if the tokens are the same size. Anthropic says the tokenizer in Claude 4.7 and later models produces about 30% more tokens for the same text [4]. Compare vendors on what a finished task costs.

I have watched context drive a bill at scale. On an AI platform at a large international marketing agency, every user's client fetched the skills library from its source, so consumption grew with users times sessions times fetches. Replacing that with one sync every ten minutes carries a forecast saving of £2 million to £4 million a year, modelled on token consumption and reviewed by the client's finance team. A later rebuild, splitting the skills so each session loads only what its task needs, cut the average cost per session by 60%, by our own measurement. The engagement file for that platform sets out both changes.

Seats look fixed until people use them

Seat prices are the easy part of the budget, with two catches. Seat tiers differ in size: a premium Claude Team seat costs five times a standard seat and carries five times the usage [2]. OpenAI's help centre likewise separates Business Standard seats from Business Premium seats, and among Business seats only Premium can spend its full allowance on GPT-6 Astra [11]. And the allowance is shared. On Claude Team and Enterprise, Claude Code draws on the same per-seat allowance as chat and Cowork [8]. On ChatGPT, Codex and ChatGPT Work share one allowance, and on credit-based agreements everyone's eligible use draws on one workspace pool [1, 5].

Coding agents are the heaviest seats. Claude Code and Codex are vendor-built versions of what we call an AI harness, and each turn carries file contents, tool calls and multi-step reasoning, which is why Anthropic tells admins to budget more for a coding seat than a chat seat [8]. OpenAI's seat, allowance and credit rules for Codex are set out in our guide to Codex pricing and credits.

Past the allowance, a seat plan becomes a metered plan. Anthropic's usage credits and OpenAI's workspace credits both let work continue, and both can be limited per person, per group or across the organisation [5, 8]. If nobody has set those limits, the seat price was never the ceiling.

Both vendors publish usage figures, and both show the same shape. Anthropic says that across enterprise deployments Claude Code averages about US$13 per developer per active day and US$150 to US$250 per developer per month, with 90% of users below US$30 per active day [8]. OpenAI's planning example for first-time enterprise ChatGPT Work adopters in Marketing puts annualised credit use at about US$53 per employee at the median, US$309 on average and US$1,154 at the 90th percentile, and calls the figures directional [5]. The average sits well above the typical user because a minority use far more. Budget for the heavy users by name, and treat any published average as a check on your own pilot, never as the plan.

Which controls stop spending, and which only warn

Every vendor offers alerts, and the documentation is explicit that they notify without stopping anything. The controls that do stop spending behave differently from one vendor to the next, and the engineering team needs to know what its application will see when one trips. The table sets out what each vendor's documentation says happens at the limit [5, 6, 7, 8].

Spending controls and what happens at the limit, from OpenAI's and Anthropic's documentation, read 29 September 2026 [5, 6, 7, 8].
DimensionWhere it appliesWhat happens at the limit
OpenAI API spend alertAn organisation or a projectA notification. API traffic continues [6]
OpenAI API hard spend limitAn organisation or a project, per monthRequests fail with a 429 error carrying a spend-limit code. Enforcement is not immediate, so recorded spend can slightly exceed the limit [6]
Claude API spend limitAn organisation or a workspace, set at or below the usage tier's monthly capRequests fail with an HTTP 400 error saying the limit was reached and when access resumes [7]
Claude API tier capThe organisation: US$500, US$1,000 or US$200,000 a month on the Start, Build and Scale tiers; the Custom tier has noneAPI use pauses until the first day of the next month unless a higher limit is granted [7]
ChatGPT usage alertsA ChatGPT workspace on a credit-based agreementA notification. OpenAI says alerts do not stop spending [5]
ChatGPT user limits and overage limitEach user, each group and the whole workspaceUser limits cap each person's eligible usage. The workspace overage limit sets how far use continues once credits run out; "No limit" is not a zero-spend cap, and nor is a group with no limit set [5]
Claude usage-credit spend limitsTeam and Enterprise: the organisation, a group or a memberRequests that would be billed to usage credits stop at the limit until an admin raises it; where the message names a plan reset time, the member can wait for that instead [8]

Two things in that table catch organisations out. The first is that one word means different things: OpenAI's API spend limit only alerts unless someone turns on the hard limit, while a Claude API spend limit stops requests once it is reached. The second is that a hard limit on a customer-facing application is an outage you scheduled yourself. I would put hard limits on experiments, internal tools and agents. On a production application I would pair a higher hard limit with alerts early enough to act on, as OpenAI suggests [6], and decide in advance what the product does when the limit trips.

Budgets, owners and chargeback

Whether you call it AI FinOps or AI spend management, the FinOps Foundation's framing is the right place to start because it is so plain: cost is price multiplied by quantity, so you manage it by lowering the rate or lowering the amount used [9]. Everything else is deciding who owns each quantity.

Give every meter one owner. Seats belong to the team the person sits in. API usage belongs to the product or use case that makes the calls, which only works if each one runs under its own OpenAI project or Anthropic workspace from the first day [6, 7]. An agent belongs to the owner of the workflow it runs, because they are the person who can say whether a run was worth what it cost.

Then decide how far to take the reporting. The FinOps Foundation separates showback, which shows each group what its scope costs, from chargeback, which posts those costs to budgets in the accounting system. It says showback is always needed, that chargeback depends on the organisation's accounting policy, and that neither is more mature than the other [10]. For most AI spend I would start with showback by owner and move to chargeback once the numbers have held steady for a few months. How the costs are booked is a question for your finance team and auditors; this is not accounting or tax advice.

Keep build money apart from run money. The Foundation's crawl, walk, run model sets cost and time limits in advance for experiments, and splits production budgets into running the system and releasing changes to it [9]. A pilot that becomes a production service without its budget moving with it is how AI spend ends up in a discretionary line with no owner at all.

A worked method, with fictional numbers

Here is the method on a made-up company. The headcounts, volumes and token counts are fictional and labelled as such; the prices are the list prices on the vendors' pages on the day we checked, so you can see how real rates behave. Replace every fictional number with one from your own pilot before a budget holder sees it.

1. Inventory the meters
List every AI line on the last three invoices and card statements, and every API key, project and workspace. Label each one seat, allowance and credits, tokens or agent run, and give it an owner.
2. Run a pilot on real work
Two to four weeks with a representative group, on the models and settings you plan to roll out. Record usage per person and per workload: the median, the 90th percentile and the highest.
3. Price each line
Seats times price. Tokens times each model's rate, split into uncached input, cached input and output. Runtime and tool calls for agents.
4. Set the limits
A hard limit on every experiment, internal tool and agent; alerts early enough to act on for production; per-person usage-credit limits on the heaviest seats.
5. Review monthly
Compare actual spend with the budget by meter and owner, and redo the arithmetic whenever a model, a price or a default changes.
A fictional 400-person company's monthly AI budget. Volumes are invented for illustration; prices are list prices in US dollars from the vendors' pages, read 29 September 2026 [1, 2, 3, 4].
DimensionFictional volumeList price usedMonthly cost
Chat seats150 peopleUS$20 a seat billed annually, the Claude Team standard and ChatGPT Business price [1, 2]US$3,000
Coding seats40 peopleUS$100 a seat billed annually, the Claude Team premium price [2]US$4,000
Usage beyond seat allowancesSix heavy users at US$120 each, from the fictional pilotUsage credits at the plan's or agreement's rateUS$720
Support assistant on the API200,000 tickets; 3,000 input and 400 output tokens each; 70% of input served from cacheUS$2 input, US$0.20 cached input and US$10 output per million tokens, the GPT-6 Sol and Claude Sonnet 5.5 price [3, 4]US$1,244, before cache-write charges
Overnight reconciliation agent900 runs; 50,000 input and 15,000 output tokens and 30 minutes of running time eachClaude Opus 5.5 at US$4 input and US$20 output per million tokens, plus US$0.08 per session-hour on Claude Managed Agents [4]US$486
TotalFive linesSeats US$7,000; metered lines US$2,450US$9,450
Monthly AI spend register
MONTHLY AI SPEND REGISTER
Month: ______   Prices checked on the vendors' pages on: ______

ONE BLOCK PER LINE OF SPEND
Line (vendor, product, project or workspace): ______
Meter: seat / allowance and credits / tokens / agent run
Owner (a named person): ______
Allocated to (team, product or workflow): ______

Budget this month: US$ ______
Limit type: hard limit / alert only / none   Limit value: US$ ______
Alert thresholds and recipients: ______

Actual this month: US$ ______   Variance: US$ ______
Usage per person or per run:
  median ______   90th percentile ______   highest ______
Unit cost (per ticket, per run, per active user): ______
Default model and settings: ______   Changed since last month? ______
Share of input tokens served from cache: ______

Action agreed: ______   By whom: ______

Reconcile against the invoice, not the dashboard: usage reports are estimates.

Seats are 74% of this fictional bill and the part nobody needs to watch. The metered 26% is where the variance lives, and two ordinary changes show why (our arithmetic on the list prices above). If a prompt change stopped the support assistant hitting its cache, that line would rise from US$1,244 to US$2,000 with no change in volume. If someone switched the same assistant to GPT-6 Astra, it would come to US$6,220, five times the line, again with no change in volume [3].

So set the limits where the variance is. In this example the seats need a quarterly headcount check and little else. The support assistant gets its own API project, with a hard limit above the pilot's 90th percentile and alerts well below it. The agent gets a workspace spend limit and a budget per run that code checks and enforces, the same way we recommend gating an agent's consequential actions. The six heavy users get per-person usage-credit limits that a manager can raise on request.

What to review every month

If the register above is kept, the monthly review should take less than an hour. These are the questions I would put to each owner.

Actual against budget, by meter and owner
Split the total. A single AI figure lets a seat saving hide a workload that is running away.
The heaviest users and workloads
Start at the top of the distribution. Check whether the heaviest use is valuable work or a session nobody cleared: Anthropic says unexpectedly high Claude Code spend on API and cloud-provider plans usually traces back to long sessions that were never cleared or to Opus left as the default model [8].
Cost per unit of work
Cost per ticket, per report or per agent run: the FinOps Foundation's cost-per-inference measure, applied to a unit the business recognises [9]. A rising total with a falling unit cost is growth. A rising unit cost is a problem.
Model mix and cache share
Which models the tokens went to, and what share of input came from the cache. Both move quietly when someone edits a prompt or a default.
Changes on the vendor side
New prices, new models and retirements. A retirement forces a model switch, and the new model's usage per task will differ, so measure again rather than carry old figures forward.
The invoice
Dashboards are estimates. OpenAI says its usage reports are planning and monitoring tools rather than invoices, and that some estimates use the overage rate rather than the committed one. Claude Code computes its cost figures at list price unless an administrator sets the organisation's contracted rates [5, 8].

Where 1AYM fits

Putting this in place is platform work as much as finance work: a project or workspace per workload, limits and alerts, per-user reporting for the coding agents, and a unit cost the business recognises. Setting up and rolling out AI work tools and coding agents for companies is core work at 1AYM; the agency platform above is the documented case, including the changes that cut its running cost. For a spend problem I would start with a fixed-scope statement of work: inventory the meters, name the owners, set the limits and measure cost per unit of work. It can start within a day of the scope being signed. If you have already scoped the job and need people to do it, we can resource it on contract from the associates who work with 1AYM, held to the same standard. Book a call below or email tayyeb@1aym.com if that would help.

For engineers: limits, error codes, attribution and caching

The implementation detail behind the controls above, for whoever owns the API projects, the gateway and the dashboards.

One project or workspace per workload
Create an OpenAI project or an Anthropic workspace for each application or agent, with its own keys, so spend, limits and reports separate at source [6, 7]. Anthropic workspaces also take their own rate limits, which cap one workload's share without touching the others; the default workspace cannot be limited [7].
Treat spend-limit errors as their own case
At a hard limit you set, OpenAI returns 429 with organization_spend_limit_exceeded or project_spend_limit_exceeded. At the usage limit OpenAI itself approves for the organisation, the code is organization_usage_limit_exceeded, and the fix is to request a higher approved limit [6]. Anthropic returns 400 invalid_request_error at a limit you set, and at the tier cap a 429 rate_limit_error with error_code enforced_spend_limit_reached and no retry-after header, where retries fail until access resumes [7]. Retrying these as rate limits wastes calls and hides the outage.
Know the tier cap
Anthropic's Start, Build and Scale tiers carry monthly caps of US$500, US$1,000 and US$200,000. At the cap, API use pauses until 00:00 UTC on the first of the next month unless the limit is raised [7]. A growing workload can reach it before anyone has set a limit of their own.
Per-user reporting for coding agents
Claude Code exports per-user token and cost metrics over OpenTelemetry on every setup; the Claude Code Analytics API and, on Enterprise, the Enterprise Analytics API return per-user reports. On Amazon Bedrock, Google Cloud and Microsoft Foundry, use OpenTelemetry or a gateway, because the analytics do not cover that usage [8]. In a Codex CLI session, /status shows the remaining limits [1].
Report at contracted rates
Claude Code's modelPricing managed setting makes /usage, the status line and OpenTelemetry report your contracted rates instead of list price. It changes the reporting, not what Anthropic charges [8].
Caching arithmetic
Anthropic charges 1.25 times the base input price to write a five-minute cache entry and 2 times for a one-hour entry, and 0.1 times to read one (0.05 times on Claude Opus 5.5), so a five-minute entry pays back after one read and a one-hour entry after two [4]. On GPT-6 Sol, cache writes are US$2.50 and cached input US$0.20 per million tokens [3].
Tool and agent overheads
Tool definitions are input tokens on every call. Anthropic's computer-use toolset adds about 4,500 input tokens a request, and fetched web content counts as input at roughly 2,500 tokens for a 10 kB page [4]. OpenAI bills eligible container sessions by the minute, with a five-minute minimum per session [3]. Claude Code agent teams use tokens roughly in proportion to the number of teammates, since each has its own context window [8].
Reasoning effort
Thinking tokens are billed as output. In Claude Code, lower the effort level for simple tasks; Opus 5.5, Sonnet 5.5 and the Fable models always use extended thinking [8].

Sources

  1. [1]OpenAI, Codex and ChatGPT Work pricing (learn.chatgpt.com), read 29 September 2026
  2. [2]Anthropic, Claude plans and pricing (claude.com), read 29 September 2026
  3. [3]OpenAI, API pricing (developers.openai.com), read 29 September 2026
  4. [4]Anthropic, Claude API pricing, including Claude Managed Agents (platform.claude.com), read 29 September 2026
  5. [5]OpenAI, ChatGPT Work: usage and cost (learn.chatgpt.com), read 29 September 2026
  6. [6]OpenAI, API spend limits (developers.openai.com), read 29 September 2026
  7. [7]Anthropic, Claude API rate limits and spend limits (platform.claude.com), read 29 September 2026
  8. [8]Anthropic, Claude Code: manage costs effectively (code.claude.com), read 29 September 2026
  9. [9]FinOps Foundation, FinOps for AI overview (finops.org), read 29 September 2026
  10. [10]FinOps Foundation, Invoicing and chargeback capability (finops.org), read 29 September 2026
  11. [11]OpenAI Help Center, ChatGPT Work and Codex (help.openai.com, checked in a browser), read 29 September 2026

AI spend questions

How much do AI tokens cost?

It depends on the model. On OpenAI's API pricing page, read 29 September 2026, GPT-6 models run from US$0.10 to US$10 per million input tokens and from US$0.50 to US$50 per million output tokens. On Anthropic's, Claude Haiku 4.5 is US$1 input and US$5 output, Claude Sonnet 5.5 US$2 and US$10, and Claude Opus 5.5 US$4 and US$20. Cached input, batch processing and fast mode each change the rate.

What is FinOps for AI?

It is the FinOps Foundation's name for managing AI cost the way FinOps manages cloud cost: tracking and reviewing AI spend and usage, setting quotas, allocating costs to owners and tying spend to business outcomes. The Foundation says the basic equation still holds, cost is price times quantity, but that AI adds new meters such as tokens, prices that change often and spend from teams outside engineering.

Do spend alerts stop AI spending?

No. OpenAI's API spend alerts and ChatGPT usage alerts notify without stopping anything. To stop spending, turn on a hard spend limit for an OpenAI organisation or project, set a spend limit on a Claude API organisation or workspace, or set user, group and overage limits in a ChatGPT or Claude workspace.

Should AI costs be charged back to departments?

Show every team what its AI use costs from the start. That is showback, and the FinOps Foundation treats it as always needed. Posting the costs to departmental budgets in the accounting system is chargeback, which depends on your accounting policy and is not a sign of maturity in itself. This is not accounting or tax advice.

How do we cut LLM costs without making the product worse?

Measure cost per unit of work first, then work down the list: route simple tasks to a cheaper model, keep prompts stable so input hits the cache, send work that can wait through a Batch API at half price, trim the context each call carries, and cap output and reasoning where the task does not need them. Test quality again after each change.

Is a seat plan cheaper than paying for the API?

For people who use AI every day in a chat or coding tool, a seat usually gives more predictable spend, because the allowance is included. For an application or an automated workload the API is the fit, billed per token. Some plans are both: Claude Enterprise is a seat price plus usage at API rates.

Further

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM