Decision guide

Private LLM: vendor API, private cloud or self-hosted

A private LLM is a large language model deployed so that your prompts, documents and outputs are handled on terms you control: not used for training, processed and stored where you chose, and kept only as long as you allow. On the vendors' own documentation, checked on 29 September 2026, there are three ways to get one. A vendor's business API from OpenAI or Anthropic is the least work and does not train on your data by default, but the vendor processes every request, and neither offers a setting that keeps inference inside the UK. A private cloud deployment on Amazon Bedrock, Microsoft Foundry or Google Cloud keeps the model's maker out of your data and can pin processing to a geography, where the model you want is offered there. Self-hosting an open-weight model puts the data and every model change under your control, and makes serving, security, evaluation and updates your job. Private settles who sees the data and where it goes. It does not make the model's answers correct or the application around it secure.

Private LLM. A large language model deployed so that the organisation using it controls who can access its prompts and outputs, where they are processed and stored, how long they are kept and whether they are used for training.

1AYM is an OpenAI Select Partner. Its founder holds personal Claude certifications. 1AYM is not an Anthropic partner.

Checked . Data handling, retention and residency terms were read on each vendor's own documentation. Vendors change prices, terms, regions, deployment types and model lists without notice, and most options apply per model rather than per platform, so confirm them on the sources below, and in your contract, before you rely on them. This page is not legal advice.

The three options side by side

What each route publishes about training, access, residency and retention, read from a UK buyer's point of view. The numbers in brackets point to the sources at the foot of the page. Where a term depends on the model or the deployment type, the cell says so rather than picking the best case.

Training, access, UK residency, retention, model changes and cost basis for a vendor API, a private cloud deployment and a self-hosted open-weight model, as published by the vendors on 29 September 2026
DimensionVendor APIPrivate cloudSelf-hosted open-weight
Examples [4, 10, 15, 18]The OpenAI API; the Claude APIAmazon Bedrock; models sold by Azure in Microsoft Foundry, which include Azure OpenAI; Google Cloud's Gemini Enterprise Agent Platform (formerly Vertex AI)An open-weight model such as gpt-oss or Llama, served on GPUs you rent or own
Used for training by default [1, 2, 10, 12, 22]No. OpenAI does not train on API data unless you opt in; Anthropic does not train on commercial products by defaultNo. AWS states Bedrock inputs and outputs are not used to train Amazon or third-party models; Microsoft states prompts are not used by model providers to improve their models; Google will not train or fine-tune on your data without permissionOnly if you fine-tune it yourself
Who can see prompts and outputs [1, 3, 5, 6, 10]The vendor. OpenAI keeps abuse-monitoring logs for up to 30 days; Anthropic keeps no conversation content by default, except 30 days for its Covered Models and up to 2 years for flagged contentThe cloud provider, not the model's maker. Bedrock and Azure both state the model provider has no access. Bedrock retains prompts for AWS review on models that require it, and Azure may sample prompts for abuse reviewYour own staff and whoever runs your infrastructure
Inference inside the UK [1, 4, 9, 11, 13]No setting for it. OpenAI stores data at rest in the UK but processes it elsewhere; Anthropic offers US or global inferencePer model: Google lists a United Kingdom (europe-west2) processing column, and its EU multi-region excludes the UK; Azure Standard deployments process in your Azure geography; Bedrock lists In-Region, Geo and Global routes per modelYes, if the hardware is in the UK
Zero data retention [1, 3, 6, 10, 12]By approval: OpenAI's Zero Data Retention or Modified Abuse Monitoring; Anthropic's zero data retention arrangementBedrock's retention mode none, unless the model requires review; Azure's modified abuse monitoring, by application; on Google, disable caching and request an abuse-monitoring exceptionSet by your own logging
When the model changes [16, 17]On the vendor's schedule: OpenAI gives at least 6 months' notice for generally available models, Anthropic at least 60 daysOn each platform's own retirement schedule, which can differ from the model maker'sWhen you decide, and every security patch is yours
How you pay [1, 4, 7, 11]Per token, with a published uplift for residency: 10% on OpenAI's residency endpoints for newer models, 1.1 times for Anthropic's US-only inferencePer token, or reserved capacity: Provisioned Throughput on Bedrock, Regional Provisioned on AzureGPU time or hardware, paid for whether or not it is busy, plus the people who run it

What private can guarantee, and what it cannot

Every one of the three routes can be made private in a useful sense. None of them is private in every sense, and the gaps are where projects get into trouble with a security reviewer late in the day.

It can keep your data out of training
All the managed routes above state that they do not train on your data by default [1, 2, 10, 12, 22]. It is now a stated term rather than a reason to pick one vendor.
It can keep the model's maker out
On Bedrock, AWS runs each provider's model in accounts the provider cannot access, so the provider never sees your prompts or Bedrock's logs [5]. Microsoft states that prompts and completions are not available to OpenAI or other providers of models sold by Azure [10].
It cannot, by itself, keep processing in one country
Storage and processing are separate settings. Google states that its endpoints do not guarantee data residency or in-region ML processing [14], and a global endpoint may process requests anywhere [13]. Azure Global deployments may be processed in any geography where the model is deployed [10]. Choose the deployment type that pins processing, and record it.
It cannot always switch off abuse monitoring
Each vendor keeps some content for abuse review in some cases: OpenAI and Google unless you are approved for an exception; Microsoft unless you are approved for modified abuse monitoring, after which automated review still runs; and Anthropic for flagged content, even under zero data retention [1, 3, 10, 12]. Ask for the exception early, because it is granted per account and sometimes per model [6].
It cannot make the application secure
A private model still answers whoever can reach it. The NCSC's guidance asks for controls on the query interface to detect attempts to extract confidential information [20]. Access control, input handling and audit sit in your application, whichever route you choose.
It cannot make the answers right
Privacy says nothing about accuracy. Evaluation does, and it is the same work on all three routes.

UK residency: where the data is stored is not where it is processed

For a UK buyer the useful question is not whether a vendor has a UK region. It is whether the model you want is processed in the UK on the deployment type you will actually use. As of the date above the vendors answer that differently.

OpenAI offers data residency in the United Kingdom for storage at rest, with regional processing not offered there; processing in-region is available in the United States, Europe (the EEA and Switzerland) and the UAE, and residency requires Modified Abuse Monitoring or Zero Data Retention [1]. Anthropic's own API offers two inference settings, US and global, and stores workspace data in the US [4].

The private cloud platforms decide per model. Google's data residency table has a United Kingdom (europe-west2) column for ML processing, model by model, and Google states that its EU multi-region excludes the United Kingdom [13]. Microsoft processes Standard and Regional Provisioned deployments within the customer-specified Azure geography, while its Data Zones are the US, the EU and Asia Pacific [11]. AWS documents, for each model, which Regions serve it In-Region, through a Geo profile or globally [9]; on a Geo profile your data stays stored in the source Region, but prompts and outputs may be processed, and stored for abuse detection, in another Region of that geography [8].

Under UK GDPR, sending personal data to, or making it accessible by, a separate organisation located outside the UK is a restricted transfer. It needs UK adequacy regulations, appropriate safeguards such as the International Data Transfer Agreement, or an Article 49 exception [21]. Whether a vendor processing your prompts abroad makes a restricted transfer depends on which entity receives the data and where it is located, so check your vendor's data processing terms. This is a reading of the ICO's guidance, not legal advice; confirm the position for your data with your own counsel.

What drives the cost

The per-token price is the number people compare and rarely the one that decides the bill. These are the drivers to model on your own traffic before choosing. We quote no figures here that the vendors do not publish.

Volume and shape of traffic
Pay-per-token routes cost what you use. Reserved capacity, such as Provisioned Throughput on Bedrock or Regional Provisioned on Azure [7, 11], and self-hosted GPUs cost the same whether traffic is heavy or idle, so they only pay back on steady, predictable volume.
Residency itself
Keeping processing in one place is priced. OpenAI charges a 10% uplift on its data residency endpoints for models released on or after 5 March 2026 [1]; Anthropic prices US-only inference at 1.1 times the standard rate [4]; AWS describes global routing as about 10% cheaper than a geographic profile [7].
Model size and hardware
For self-hosting, the model you pick sets the hardware. OpenAI describes gpt-oss-120b as fitting on a single 80GB GPU and gpt-oss-20b as running within 16GB of memory [18]; larger models need more, and a second copy for resilience needs its own.
People
A self-hosted model needs someone on call for serving, scaling, patching and incidents. Put that salary in the comparison next to the hardware.
The design around the model
How the application calls the model moves the bill as well as where it runs. On the AI platform of a large international marketing agency, average cost per session fell 60% after we re-architected its skills estate, by our own measurement, as set out in engagement file D-01.

The operations burden: what LLMOps means in practice

LLMOps is the work of keeping a language-model system correct, affordable and accountable after launch. The managed routes take some of it off you. None takes all of it, and self-hosting hands you the rest.

Evaluation
A regression suite that runs whenever a prompt, a model or a retrieval source changes, so a quality drop is caught before users see it. On one production estate we run around 1,170 automated tests on the speaking pipeline, including a suite where disabling any one guardrail breaks specific frozen cases.
Monitoring
Quality, latency, cost and where each request actually ran. The vendors expose the last of these: Bedrock logs the processing Region in CloudTrail [7], and Anthropic's API returns the inference geography with each response [4].
Model updates
Managed models retire on the vendor's schedule, at least 6 months' notice for OpenAI's generally available models and at least 60 days for Anthropic's [16, 17], and each migration needs the regression suite. Self-hosted models change only when you change them, and you verify the files: the NCSC recommends cryptographic hashes or signatures of model files [20].
Fallbacks
A second model or provider for when the first is unavailable or retired. In production today we run OpenAI audio models for speech assessment with a Gemini fallback. Fallbacks cover availability, not judgement. Judgement is the job of the agent-proposes, verifier-gates pattern, which holds that a model should never be the only gate on a consequential action: deterministic validators and human review gates decide what proceeds.
Licences
Open-weight does not mean unconditional. gpt-oss is released under Apache 2.0 [18]; the Llama 3.1 licence requires a separate licence from Meta for licensees above 700 million monthly active users and a visible "Built with Llama" notice when you distribute it [19]. Read the licence of the exact model you deploy.
Incidents
The NCSC's guidance treats security incidents affecting AI systems as inevitable and asks for response, escalation and remediation plans that cover them [20]. On a managed route the vendor runs part of that plan; self-hosted, all of it is yours.

Which option fits which situation

Find the constraint that binds hardest. It usually decides the route on its own, and the others become tuning.

You need a frontier model and standard business terms are enough
A vendor API, with an application for Zero Data Retention or Modified Abuse Monitoring if prompts carry sensitive data [1, 3]. It is the fastest route and the one with the least to run.
Processing must stay in the UK
As of the date above, neither OpenAI nor Anthropic offers a setting on its own API that keeps inference in the UK [1, 4]. Use a private cloud deployment for a model its provider lists for UK processing, or self-host in the UK. Check the model, the region and the deployment type together.
The model's maker must never see the data
A private cloud deployment, where the provider states the model's maker has no access [5, 10], or self-hosting.
The model must not change without your say-so
Self-host, or pin a model version on a managed route and budget a tested migration before each retirement date [16, 17].
Volume is high, steady and predictable
Reserved capacity or self-hosting may cost less than tokens. Model it on a month of your real traffic before committing, including the people.
Nobody on the team runs GPUs today
Do not start with self-hosting. A private cloud deployment gives most of the control without the serving and patching.
You are not yet sure a build is the answer
If an off-the-shelf assistant would do, our side-by-side of the ChatGPT and Claude enterprise plans sets out what each publishes on price, controls and data location. If the question is which build, if any, pays back, that is what our AI strategy work settles first.

Where 1AYM stands

1AYM is an OpenAI Select Partner, and the Claude certifications on this site belong to our founder personally, not to the company. We choose models per workload across OpenAI, Anthropic, Gemini and open-source components.

The residency decision we have made on our own work is a database one: the production estate of a government-accredited EdTech runs its database on Postgres in Google Cloud's Doha region, me-central1, to meet Gulf data-residency requirements, and where its models run was kept a separate decision. We build production AI systems designed around who may see what, what happens when a model fails, what gets logged and where a human signs off, and hand them over with evaluation suites, audit logging and a runbook.

For engineers: residency and retention controls to check

The settings behind the table, each with its source. Most are per project, per Region or per model, so check them where your traffic actually runs.

OpenAI data residency
A project setting, available by eligibility through OpenAI's sales team, with a regional domain per region; the UK prefix is gb.api.openai.com, storage only. It needs Modified Abuse Monitoring or Zero Data Retention, and does not cover system data such as account metadata [1].
Anthropic inference_geo
Set per request or as a workspace default, with allowed_inference_geos to restrict it. Values are us and global, on Claude 4.6 and later models; the response's usage.inference_geo field records where inference ran. Workspace geo is us only and cannot be changed after creation [4].
Bedrock data retention mode
none, default or aws_review, set per Region at account or project level; the setting does not propagate to other Regions. A Service Control Policy can deny any mode other than none. Models that require review become unavailable under none [6].
Bedrock routing evidence
Geo and Global inference profiles route across Regions; CloudTrail records the processing Region in additionalEventData.inferenceRegion in your source Region [7]. Regions an SCP blocks break the profile unless you add an inference-profile exception [8].
Azure modified abuse monitoring
Granted on application. Once approved, the resource's ContentLogging capability reads false in the portal's JSON view or the Azure CLI, which is the evidence to file [10].
Google Cloud retention switches
In-memory caching of inputs and outputs, with a 24-hour TTL, is on by default and can be disabled per project. Request-response logging is off by default. The Interactions API stores data unless store is false. Grounding with Google Search keeps logs for up to three days and cannot be switched off [12].
Self-hosted model files
Verify model weights against published hashes or signatures before deployment, protect the weights as you would credentials, and put the serving endpoint behind your own authentication and query controls [20].

Sources

  1. [1]OpenAI, Data controls in the OpenAI platform (training, abuse monitoring, data residency), read 29 September 2026
  2. [2]Anthropic Privacy Center, Is my data used for model training? (dated 18 August 2026), read 29 September 2026
  3. [3]Anthropic, API and data retention, read 29 September 2026
  4. [4]Anthropic, Data residency (inference geo and workspace geo), read 29 September 2026
  5. [5]AWS, Amazon Bedrock data protection, read 29 September 2026
  6. [6]AWS, Amazon Bedrock data retention, read 29 September 2026
  7. [7]AWS, Amazon Bedrock cross-Region inference, read 29 September 2026
  8. [8]AWS, Amazon Bedrock geographic cross-Region inference, read 29 September 2026
  9. [9]AWS, Amazon Bedrock supported Regions and models for inference profiles, read 29 September 2026
  10. [10]Microsoft, Data, privacy and security for Foundry Models sold by Azure, read 29 September 2026
  11. [11]Microsoft, Foundry deployment types and data processing location, read 29 September 2026
  12. [12]Google Cloud, Gemini Enterprise Agent Platform and zero data retention (last updated 28 September 2026), read 29 September 2026
  13. [13]Google Cloud, Gemini Enterprise Agent Platform data residency (last updated 28 September 2026), read 29 September 2026
  14. [14]Google Cloud, Gemini Enterprise Agent Platform locations (last updated 28 September 2026), read 29 September 2026
  15. [15]Google Cloud, Gemini Enterprise Agent Platform (formerly Vertex AI), read 29 September 2026
  16. [16]OpenAI, Deprecations (notice periods), read 29 September 2026
  17. [17]Anthropic, Model deprecations (notice periods), read 29 September 2026
  18. [18]OpenAI, gpt-oss-120b model card (licence and hardware), read 29 September 2026
  19. [19]Meta, Llama 3.1 Community License Agreement, read 29 September 2026
  20. [20]NCSC, Guidelines for secure AI system development: secure deployment (27 November 2023), read 29 September 2026
  21. [21]ICO, International transfers: a guide (last updated 15 January 2026), read 29 September 2026
  22. [22]AWS, Amazon Bedrock FAQs (training on inputs and outputs), read 29 September 2026

Private LLM questions

What is a private LLM?

A large language model deployed so that your organisation controls who can see its prompts and outputs, where they are processed and stored, how long they are kept and whether they are used for training. It can be a vendor's API under business terms, a model run in your own cloud account, or an open-weight model you host yourself.

Is a private LLM the same as a self-hosted LLM?

No. Self-hosting is one way to get a private LLM, and the one with the most control and the most work. A private cloud deployment on Amazon Bedrock, Microsoft Foundry or Google Cloud also keeps the model's maker out of your data, without running the serving yourself.

Does OpenAI or Anthropic train on our API data?

Not by default, on their own terms read on 29 September 2026. OpenAI states that API data is not used to train its models unless you opt in. Anthropic states that it does not train on inputs or outputs from its commercial products by default, with exceptions such as feedback you choose to send.

Can we keep inference in the UK?

Not as a setting on OpenAI's or Anthropic's own API, as read on 29 September 2026. OpenAI offers UK storage at rest with processing elsewhere, and Anthropic offers US or global inference. Google Cloud lists UK (europe-west2) processing per model, Azure processes Standard deployments within your Azure geography, and AWS lists per model which Regions serve it in-Region. Self-hosting in the UK keeps everything there. Check the exact model and deployment type.

Is Azure OpenAI private?

Microsoft states that prompts and completions sent to models sold by Azure, including Azure OpenAI models, are not available to OpenAI, are not used by model providers to improve their models, and are not available to other customers. Where they are processed depends on the deployment type: Standard deployments stay in your Azure geography, Data Zone deployments in the data zone, and Global deployments may be processed anywhere the model is deployed.

What is LLMOps?

The operational work of running a language-model system after launch: evaluation when prompts or models change, monitoring of quality, cost and where requests ran, migrations when models retire, fallbacks, licences and incident response. Managed routes carry part of it; self-hosting carries none of it for you.

How much does a private LLM cost?

It depends on four drivers more than on the per-token price: how steady your volume is, whether you pay for residency, which model size you need, and who runs it. Pay-per-token routes cost what you use; reserved capacity and self-hosted GPUs cost the same busy or idle. The vendors publish residency uplifts of 10% on OpenAI's residency endpoints for newer models and 1.1 times for Anthropic's US-only inference.

Further

  • Production AI systems · The service behind this guide: copilots, workflow agents and LLM middleware built for production, with evaluation suites, audit logging and human review.
  • Gulf AI data residency · The same storage-versus-processing question for Saudi Arabia, the UAE and Qatar, with the Gulf law that applies.
  • AI governance framework · Where the hosting decision is recorded as a control with an owner and evidence.

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM