Template
AI project statement of work: a template and checklist
A statement of work (SOW) for an AI project sets out the objective, the deliverables and how each one will be accepted, plus the terms an ordinary software SOW leaves out: which of your data the supplier may use and for what, how the system is tested and on which cases, what happens when the underlying model or provider changes, where a person reviews the output, who owns the code, prompts and test sets, and what is handed over at the end. It usually sits under a master services agreement (MSA), which holds the legal terms, while the SOW describes the project. The section that matters most is acceptance: write it as tests on real cases, with pass thresholds agreed before the work starts.
AI statement of work. The project document, usually signed under a master services agreement, that fixes what an AI supplier will deliver, how it will be tested and accepted, which data it may use, and what the buyer owns and receives at the end.
Checked . The federal regulations, OMB guidance, NIST framework and copyright law this page cites were read on each publisher's own site on this date. Law and guidance change, so check the current text before you rely on it. This template is not legal advice: have your own counsel review any contract before you sign it.
MSA vs SOW: which terms go where
A services contract between two companies usually comes in two layers. The master services agreement (MSA) holds the legal terms that apply to every piece of work: liability, confidentiality, insurance, payment terms and termination. Each statement of work (SOW) then describes one project under it: what gets built, by when, for how much, and how you will decide it is done. You negotiate the MSA once and sign a new SOW for each project.
The AI-specific terms can live in either, and I would split them by how often they change. Rules that hold for every project, such as whether the supplier may train models on your data, belong in the MSA or a data addendum to it. Anything that depends on this project, such as the test cases, the pass thresholds and the model in use, belongs in the SOW. If you have no MSA yet, the template below still works as a project document, but your counsel will need to put the legal terms around it.
| Dimension | Master services agreement (MSA) | Statement of work (SOW) |
|---|---|---|
| What it is for | The legal terms for the whole relationship | One project: what is delivered, when, for how much, and how it is accepted |
| How often it is signed | Once, and amended rarely | Once for each project, with signed change orders during it |
| Typical contents | Liability, indemnities, confidentiality, insurance, payment terms, termination and governing law | Objective, scope, deliverables, milestones, acceptance tests, fees, named people and handover |
| AI terms that belong here | Standing rules: use of your data for training, data security and retention, ownership of work product, disclosure of the supplier's own AI use | Project terms: the data sources, the models and providers in use, the test set and thresholds, human review points and the handover list |
| If the two conflict | An order-of-precedence clause says which document wins | Name, explicitly, any MSA clause this SOW is meant to override |
What an AI statement of work needs that a software SOW does not
A software SOW assumes that once a feature is specified, you can test whether it works: the form either saves the record or it does not. An AI system gives you a spread of results instead. It will be right on many cases and wrong on some, and the mix moves when the data, the prompt or the model changes. That changes three parts of the document.
Acceptance becomes a measurement. You agree a set of real cases and a pass rate before the work starts, and the system passes or fails against them. US federal buying has asked for this shape for a long time: the Federal Acquisition Regulation (FAR) tells agencies, where practicable, to describe work by the results required rather than by how it is done or the hours spent, and to assess it against measurable performance standards [1]. A 2025 memorandum from the White House Office of Management and Budget (OMB) on how federal agencies buy AI goes further. It tells agencies to run independent evaluations on data the vendor cannot access, and encourages them to require vendors to meet performance standards before deploying a new version, or to roll back to the previous one [2].
Your data needs its own terms. The supplier needs it to build and test the system, so the SOW should say which data, through what access, for what use, for how long, and what happens to it at the end. The same memorandum tells agencies to write into their contracts a permanent prohibition on using their non-public data to train publicly or commercially available AI, absent explicit consent [2]. A private company can ask for the same.
And the system depends on a model that someone else controls. Providers release new versions and retire old ones on their own timetable, so the SOW should name the model and provider in use, say who is told before either changes, and treat a change as something that reruns the acceptance tests.
The template, section by section
Copy it into your own document, delete what does not apply and fill in the brackets. The sections after it explain the parts that decide most disputes. It is a starting point for your counsel, not a finished contract.
AI PROJECT STATEMENT OF WORK SOW number: [number] Under master services agreement: [name and date of the MSA] Customer: [legal name] Supplier: [legal name] Effective date: [date] Order of precedence: if this SOW conflicts with the MSA, the MSA governs, except for these clauses, which this SOW overrides: [list, or none]. 1. OBJECTIVE The business result this project exists to produce: [one or two sentences]. Baseline today: [time, error rate or cost per case], measured on [date] by [method]. 2. SCOPE In scope: [workflows, systems, user groups]. Out of scope: [what the Supplier will not do, including any adjacent workflow a reader might assume is included]. Assumptions: [what both parties are relying on, such as data quality or system access]. Customer responsibilities: [access, data, named people, review time each week, approvals], each with a date. 3. DELIVERABLES For each deliverable: name, description, format, and where it will live (the Customer's repositories and accounts). D1. [Working system for workflow X, running in environment Y] D2. [Evaluation suite and acceptance test set] D3. [Documentation: architecture overview, runbook, decision log, documentation of data, models and prompts] D4. [Knowledge transfer sessions] 4. MILESTONES AND SCHEDULE M1. [Description], due [date], deliverables [D1, D2], payment [amount or share of the fee] on acceptance. M2. [Repeat for each milestone.] 5. ACCEPTANCE Acceptance test set: [number] real cases chosen by [Customer role], with the correct outcome agreed for each. The Customer holds the set. [Number or share] of the cases are withheld from the Supplier until acceptance testing. Pass criteria: [task measure] of at least [value] on the acceptance test set, and [service measure, such as time per case or reviewer override rate] of [value]. Test procedure: run by [who], in [environment], within [number] business days of delivery. Acceptance period: the Customer accepts or rejects in writing within [number] business days of the test run, giving reasons for any rejection. On failure: the Supplier fixes and resubmits within [number] business days. After [number] failed rounds: [remedy, such as a fee reduction, re-scoping or termination of this SOW]. 6. EVALUATION AFTER ACCEPTANCE The evaluation suite runs on every change to a prompt, model, tool or rule, and results are recorded in [location]. Scores that block a release: [list, with thresholds]. Review of live performance: [frequency], reported to [role]. 7. DATA Data sources the Supplier may access: [list, each with its access method and environment]. Permitted use: only to perform this SOW. Customer data, prompts and outputs are not used to train or improve any model or product outside this SOW without the Customer's prior written consent. Storage and processing: [where data is stored and processed, including each model provider that receives it]. Retention and deletion: [period]. Deletion confirmed in writing at the end of this SOW. Regulated data: [personal data, protected health information, customer financial information: yes or no for each]. If yes, the agreement it requires: [name the agreement], signed before any access. 8. MODELS, PROVIDERS AND CHANGES Models and providers in use: [names and versions]. Accounts and keys: held in the Customer's name. Changes: the Supplier gives [number] business days' notice before changing a model, provider or AI feature, reruns the acceptance tests before the change goes live, and can roll back to the previous version. AI tools used by the Supplier to produce the deliverables: [list, or to be listed in the handover pack]. 9. HUMAN REVIEW AND CONTROLS Actions that need a named person's approval before they take effect: [list]. Logging: [what is logged, where, and who can read it]. Switch-off: [how the capability is turned off, and who can do it]. 10. INTELLECTUAL PROPERTY Deliverables: source code, prompts, configuration, evaluation sets, infrastructure definitions and documentation produced under this SOW. The Supplier assigns its rights in them to the Customer, in writing, on [payment or acceptance]. Supplier's pre-existing materials: [list]. Licence to the Customer: [scope]. Third-party and open-source components: listed in the handover pack with their licences. 11. CHANGE CONTROL Either party may propose a change in writing. The Supplier states its effect on scope, fees, schedule and acceptance criteria within [number] business days. No change takes effect until both parties sign a change order. 12. ROLES AND GOVERNANCE Customer sponsor: [name, role]. Customer process owner: [name, role]. Supplier lead: [name, role]. Key people: [names], replaced only with the Customer's agreement. Progress review: [frequency], with a written note of decisions. Escalation: [route]. 13. FEES AND PAYMENT Fees: [fixed fee, or time-based with a cap], under the MSA's payment terms. Payments are tied to the milestones in section 4 and fall due on acceptance. Expenses: [policy, or none]. 14. HANDOVER AND EXIT Handover pack: runbook, architecture overview, decision log, documentation of data, models and prompts, the evaluation suite, and the list of third-party components. Knowledge transfer: [number] sessions with [team]. Handover test: the Customer's engineers deploy one change without the Supplier's help. On completion or termination: Supplier access revoked, Customer data returned or deleted, deletion confirmed in writing. 15. TERMINATION Either party may end this SOW as the MSA provides. On early termination the Customer pays for accepted milestones and for work in progress to [date], and receives every deliverable in its current state. SIGNATURES Customer: [name, title, date] Supplier: [name, title, date]
Scope and deliverables: write down the result
Scope sections fail in a predictable way. They list activities, such as design, build and test an AI assistant, instead of results, so a supplier can do all of them and still hand over nothing you can use. Write the objective as the change you want in a piece of work you can measure, with today's baseline beside it, so everyone can see whether it moved. It is the rule the FAR sets for performance-based federal contracts: describe the results required, not the method or the hours [1].
Be just as specific about what is out of scope. Adjacent workflows, extra user groups, other languages and integrations with systems nobody named are where the arguments start. The customer responsibilities matter as much, because access, data and a process owner's time set the pace of an AI project far more often than the engineering does. A date against each one makes a delay visible, and shows whose it was.
List deliverables as things you can open or run: a working system in a named environment, a test suite, a runbook. Say that they live in your repositories and your accounts from the first commit, because moving them at the end is slow and tends to go badly.
Acceptance and evaluation: agree the test before the result
This is the section I would spend the most time on. An AI deliverable is accepted against a test set: real cases, chosen by your process owner, with the right outcome already agreed for each. Set the pass threshold before anyone sees a score. Set it afterwards and it tends to land wherever the scores did.
Keep part of the test set back. OMB tells federal agencies to run independent evaluations on data the vendor cannot access, as close as possible to the data the system will meet in use [2]. The logic holds for a private buyer too: if the supplier has tuned against every case you hold, a pass tells you less than it appears to. The AI Risk Management Framework from the National Institute of Standards and Technology (NIST) asks for test sets, metrics and the tools used in testing to be documented [3], which is what lets someone rerun the result after the supplier has gone.
Measure two things. One is how well the system does its task. The other is whether the work improved: time per case, error rate, or how often reviewers override the output. A system can score well on the first and still be ignored by the people meant to use it. An acceptance line can be as plain as this, with your own numbers in the brackets: the system proposes the correct [field] for at least [X]% of the [N] cases in the acceptance test set, and sends every case it is less sure of than [threshold] to a reviewer.
Then write down what happens on a fail: a fix and a retest, a limit on the number of rounds, and a remedy after the last one. Without that, a failed test turns into a negotiation. For actions with consequences, put the approval step in the SOW as well; the pattern where the model proposes and a deterministic check decides is one way to build it.
Data, models and the changes you do not control
Name every data source, how the supplier reaches it and in which environment. Limit use to the project, and say in plain words that your data, prompts and outputs will not be used to train or improve any model or product outside it without your written consent: the term OMB tells federal agencies to write in [2]. Set a retention period, and ask for deletion to be confirmed in writing at the end.
Regulated data brings its own agreement. If your company is a covered entity under the Health Insurance Portability and Accountability Act (HIPAA) and the supplier will create, receive, maintain or transmit protected health information on your behalf, HIPAA requires its assurances to be documented in a written contract or other written arrangement that meets the business associate requirements [8]. If your company is a financial institution within the jurisdiction of the Federal Trade Commission (FTC), the FTC's Safeguards Rule requires you to choose service providers able to protect customer information, to require those safeguards by contract, and to assess the providers periodically [7]. Neither agreement belongs inside a SOW. The SOW's job is to flag the data and make signing the right agreement a condition of access. This is not legal advice.
For models, name the model and provider in use, and keep the accounts and keys in your company's name. A provider can change or retire a model version on its own schedule, so treat a model change as a release: notice first, the acceptance tests rerun before it goes live, and a way back to the previous version. OMB's memorandum lists the same three as terms for federal agencies to consider: notice before new AI features are added, performance standards met before a new version is deployed, and a roll-back if it falls short [2].
Intellectual property: code, prompts and test sets
Under US copyright law, copyright belongs first to the author of a work, and a transfer of ownership is only valid in writing, signed by the owner or their authorised agent [6]. A supplier's code is the supplier's until they sign it over. The work-made-for-hire label does less than many buyers expect: for work by someone who is not your employee, it covers only nine listed kinds of commissioned work, and only with a signed written agreement, and software is not named among them [5]. So ask for an express written assignment, and list what it covers: source code, prompts, configuration, evaluation sets, infrastructure definitions and documentation, not only software.
Material generated by AI adds a complication. The US Copyright Office concluded in January 2025 that copyright does not extend to purely AI-generated material, or to material where there is not enough human control over the expressive elements, and that prompts alone do not currently provide that control [4]. Work made with AI help is still protected where a person's contribution is enough, which the Office says is decided case by case [4]. The practical consequence for a SOW: an assignment passes on whatever rights exist. So also ask the supplier to say which AI tools it used, to list third-party and open-source components with their licences, and to confirm that you may use, change and run everything delivered. NIST's framework asks organisations to have policies for AI risks from third parties, including the risk of infringing someone else's intellectual property [3]. This is not legal advice.
Change control and handover
AI projects change scope more than most, because the first results teach you something about the problem. That is healthy if it goes through a change order: a written request, the supplier's statement of what it does to fees, schedule and acceptance criteria, and both signatures before the work starts. A change agreed on a call and never written down is how a fixed-scope project quietly becomes an open-ended one.
Write the handover as a test. The pack holds a runbook, an architecture overview, a dated decision log, documentation of the data, models and prompts, the evaluation suite and the list of third-party components. Then one of your engineers deploys a change without the supplier's help. OMB counts knowledge transfer, data and model portability, and rights to the code and models produced under a contract among the protections against vendor lock-in [2]. If your own team cannot ship that change, the handover has not happened, whatever the pack says.
For IT operations work, write the order of handover into the scope as well: the agent suggests before it acts, and fixes that change systems come later, once they have passed review. Our guide to AI for IT operations sets out that order.
Checks before you sign
Read the draft against these ten lines. Any line the SOW does not answer is a question for the supplier before signature, not after.
- Objective
- A measured baseline sits beside the objective, so anyone can tell whether it moved.
- Out of scope
- Written down, including the adjacent workflows a reader might assume are included.
- Acceptance tests
- Real cases with agreed outcomes, thresholds set now, and part of the set held back until acceptance.
- Failed tests
- A fix and retest, a limit on rounds, and a remedy after the last one.
- Data use
- Project use only, no training without your written consent, and deletion confirmed in writing.
- Regulated data
- Flagged, with the right agreement signed before any access.
- Model changes
- Notice before a change, the acceptance tests rerun, and a way back to the previous version.
- Ownership
- A written assignment covering code, prompts, test sets and configuration, and a list of third-party components.
- Change orders
- Signed by both parties before the changed work starts.
- Handover
- Your own engineers ship a change without the supplier's help.
Where 1AYM fits
A fixed-scope engagement at 1AYM is written as a statement of work with defined outputs, a defined price and acceptance criteria, much like the template above. We take small fixed-scope statements of work as well as larger builds, and work can start within a day of the scope being signed. Where the contract needs it, copyright in the deliverables is assigned to you, and US clients can contract through 1AYM's US entity.
If you cannot yet write the objective or the acceptance tests, that is what our AI opportunity and feasibility sprint is for: typically two to four weeks, ending in a scope you can put in a SOW. A fixed-scope architecture and production build then delivers it. Depending on the scope, the SOW names the evaluation suite, human review points and handover pack as deliverables. If you already have a scoped job and need people to deliver it, we can resource it with embedded engineers on contract, drawn from the collective of associates who work with 1AYM and held to the same standard. On the 30-minute call, your account manager will go through your draft SOW, from any supplier. That is a working read, not a legal review.
For engineers: making the SOW testable
The terms above only hold if someone can check them in the repository, the pipeline and the accounts. These are the specifics I would want written into the SOW or its handover pack.
- Acceptance harness
- The acceptance test set, expected outcomes and thresholds are versioned in the customer's repository, and the harness that scores them runs in the customer's CI. Acceptance is a passing run on a tagged commit.
- Held-out cases
- A slice of the test set sits outside anything the supplier can read, and runs only at acceptance, so a pass cannot come from tuning against the answers.
- Model pinning
- Models are named by dated version or snapshot where the provider offers one, not by a floating alias, and the version is recorded in every log line, so a provider-side change shows up as a version change the suite runs against.
- Release gates
- Any change to a prompt, model, tool definition or rule runs the evaluation suite, and a score below its threshold fails the build. The previous version stays deployable, so a rollback is a configuration change.
- Accounts and keys
- Model provider accounts, API keys and cloud projects are in the customer's name. Supplier engineers use individual identities through the customer's identity provider, revoked at exit in one step.
- Data flow register
- One table lists each data source, its access path and environment, every model provider that receives it, retention at each hop, and whether the data is regulated. The SOW's data section points to it.
- Logs
- One structured record per action: input reference, model and prompt version, check results, approver and timestamp, readable only by the roles that need it.
- Provenance
- A list of third-party and open-source dependencies with their licences, generated from the build rather than typed by hand, and a note of the AI coding tools used on the deliverables.
- Handover run
- The handover test is a real change: a customer engineer makes it, the pipeline tests and deploys it, and nobody from the supplier touches it.
Sources
- [1]Federal Acquisition Regulation 37.602, Performance work statement (FAC 2026-01, effective 13 March 2026), read 29 September 2026
- [2]OMB Memorandum M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government (3 April 2025), read 29 September 2026
- [3]NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023), read 29 September 2026
- [4]US Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability (29 January 2025), read 29 September 2026
- [5]US Copyright Act, 17 U.S.C. section 101: definition of a work made for hire, read 29 September 2026
- [6]US Copyright Act, 17 U.S.C. sections 201 and 204: ownership of copyright and transfers in writing, read 29 September 2026
- [7]FTC Standards for Safeguarding Customer Information, 16 CFR 314.1 and 314.4 (eCFR, up to date as of 25 September 2026), read 29 September 2026
- [8]HIPAA Privacy Rule, 45 CFR 164.502(e): business associate disclosures (eCFR, up to date as of 25 September 2026), read 29 September 2026
Statement of work questions
What is the difference between an MSA and a SOW?
A master services agreement holds the legal terms for the whole relationship, such as liability, confidentiality, payment terms and termination, and is signed once. A statement of work sits under it and describes one project: scope, deliverables, schedule, fees, acceptance and handover. For AI work, standing rules on the use of your data belong in the MSA, and the project's test cases and models belong in the SOW.
Who should write the statement of work, the buyer or the supplier?
Either can draft it, and suppliers often do. Whoever drafts it, the buyer should write the objective, the out-of-scope list and the acceptance tests, because those decide what the buyer is paying for. A supplier's draft that leaves acceptance vague is the first thing to push back on.
What acceptance criteria should an AI project use?
A pass rate on a set of real cases whose correct outcomes your own people have already agreed, set before work starts, plus a measure of whether the work improved, such as time per case or how often reviewers override the output. Hold part of the test set back until acceptance, and write down what happens after a failed test.
Can we own the copyright in AI-generated code?
The US Copyright Office concluded in January 2025 that copyright does not extend to purely AI-generated material, and that prompts alone do not give enough control; whether a person's contribution is enough is decided case by case. So have the supplier assign whatever rights it holds in writing, say which AI tools it used and confirm you may use everything delivered. Based on the Office's report and Title 17, read on 29 September 2026; not legal advice.
Should an AI project be fixed price?
When the output and its acceptance tests can be written down, a fixed price against them is usually the cleaner buy. When they cannot yet, pay for a short fixed-scope discovery that produces them, then price the build. Put a price on a vague output and the supplier either pads it to cover what is unknown or gets it wrong.
Do we need a business associate agreement with an AI supplier?
If you are a HIPAA covered entity and the supplier will create, receive, maintain or transmit protected health information on your behalf, HIPAA requires its assurances to be documented in a written contract or other written arrangement that meets the business associate requirements. Make signing it a condition of data access in the SOW. Based on 45 CFR 164.502(e), read on 29 September 2026; not legal advice.
Further
- AI strategy and roadmap · The fixed-scope sprint that turns an idea into an objective and acceptance tests you can put in a SOW.
- Production AI systems · The fixed-scope architecture and build, with evaluation, review points and handover designed in.
- AI implementation: the first 30 days · What the first month under a signed SOW should produce, week by week.
- Hiring an AI engineer · When a fixed-scope SOW is the right buy, and when a hire or a contract is.
- Embedded Engineers on Contract · For a job that is already scoped: engineers inside your programme, held to the same standard.
- Choosing an AI supplier · Four checks to run on any supplier before you sign anything with them.
We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.
Last reviewed · 1AYM