Guide

Adding AI to your software product: partner, hire or both

A software company adding AI features to its product has four routes: build with the engineers it already has, hire engineers with production AI experience, bring in a partner to build the feature, or have a partner build alongside its own team. The right one depends on how central AI will be to what customers pay for, how soon they need it, and whether anyone in-house can own evaluation: the automated tests that show whether a model's output is good enough to ship. Whichever route you take, your own team should end up owning the feature, its evaluation set and its running cost. A partner that cannot hand those over is renting you a feature.

AI feature. A capability in a software product whose output is generated, ranked or classified by a machine learning model, usually a large language model called through a vendor's API, rather than by rules written in code.

1AYM is an OpenAI Select Partner. Its founder holds personal Claude certifications. 1AYM is not an Anthropic partner.

Checked . US statutes, US Copyright Office and FTC publications, NIST and OWASP guidance, and the OpenAI and Anthropic evaluation guides were read on the publishers’ own pages. Vendor documentation changes without notice, so confirm the details on the sources below before you rely on them. Nothing here is legal advice.

Why the demo is the easy part

A working demo of an AI feature takes an engineer an afternoon: send the customer's text to a model and show what comes back. That speed is why AI items now sit on so many roadmaps, and it is also what throws the planning off, because the demo skips every part that decides whether the feature can ship.

The model does not reliably give the same answer twice. OpenAI's own evaluation guide says models sometimes produce different output from the same input, which makes traditional software testing insufficient. NIST's profile for generative AI names the failure directly: confabulation, the production of confidently stated but erroneous or false content. So an AI feature needs four things a normal feature does not. It needs test cases that measure whether the output is good enough, limits on what the model can see and do, a cap on what each customer's usage can cost, and a record that lets support explain an answer a customer disputes.

Those four pieces are where the time goes, and they are what to plan the staffing around, because whoever owns them owns the feature. A feature that takes actions for a user, rather than drafting text for them to read, needs more controls again, and our guide to where AI agents fit in a business sets those out.

Four ways to staff it, side by side

All four routes can ship an AI feature. They differ in when the work starts, who carries the risk and what you are left owning. These are typical patterns, not rules, so ask any candidate or supplier where they differ.

Four ways a software company can staff its first AI features: what each costs, when it starts, its main risk, who builds and supports the feature, and when it fits
DimensionYour current teamNew hiresA partner builds itA partner alongside your team
What you pay forEngineering time taken from the existing roadmapSalaries and recruiting, before any feature existsA defined output, or a team's time, from an outside firmThe partner's time plus your engineers' time, with teaching built into the plan
When work startsNow, if anyone has capacityAfter the hiring round: search, offers, notice periods and ramp-upWhen the scope is signedWhen the scope is signed
Main riskEvaluation and guardrails are learned on live customersHiring before you know what the feature needs, with nobody in-house able to judge the candidatesThe feature works, but only the partner can change itThe handover slips because your engineers never get the hours
Who builds the evaluation setYour team, often lateThe new hires, once they arriveThe partner, and it is yours only if the contract says soBoth, with your team owning it by the end
Who supports it after launchYour teamYour teamThe partner under a support agreement, or your team after a handoverYour team
Fits whenThe first feature is narrow and a person reads the output before acting on itAI will be a lasting part of what customers pay for and the role is already clearCustomers need the feature sooner than a hiring round allows and the scope can be written downAI will be central to the product and you want your own team running it

Building with the team you already have

If your engineers have capacity and someone senior wants to own evaluation, this is the cheapest route in cash and the most expensive in lessons. The model APIs are well documented, and your team already knows the product, the data and the customers, which is the knowledge an outsider takes longest to pick up.

The catch is that evaluation, prompt injection and cost control are new disciplines for a team that has built deterministic software. Learned on the job, they are learned on live customers. This route works best for a narrow first feature, such as a summary or a drafting aid that a person reads before acting on it, where a weak answer is visible and cheap to fix.

Hiring for it

Hiring makes sense when AI will be a lasting part of what customers pay for and you know what the role will spend its first year doing. A permanent engineer with production AI experience keeps the knowledge inside the company, which no partner can do.

Two problems come with it. The hiring round comes before any feature exists, and the roadmap waits through search, offers, notice periods and ramp-up. And the first AI hire is hard to judge when nobody in the company has shipped an AI feature, so you are assessing experience you cannot yet check. My preference is to ship the first feature, then hire into a role the company now understands. 1AYM is one of the partners that could build that first feature, so weigh my preference accordingly.

If you bring in independent contractors rather than employees, check who owns the code. The US Copyright Office's Circular 30 explains that work an employee creates as part of their regular duties is a work made for hire, owned by the employer. Commissioned work is made for hire only if it falls within nine categories the statute lists and both parties sign a written agreement saying so. Software is not named among those categories, so the dependable route is a written assignment of copyright, which 17 U.S.C. § 204(a) says is not valid unless it is in writing and signed by the owner of the rights.

Bringing in a partner

A partner that has shipped AI features before brings the missing disciplines on the first day: the evaluation harness, the guardrails and the cost controls. The work starts when the scope is signed rather than when a hiring round closes, and a fixed-scope contract prices a defined output, which moves the risk of overrun off your side of the table.

The risk is ending up renting a feature. If the code sits in the partner's repositories, the model keys in the partner's accounts and the evaluation set in someone's head, you have a working feature that only the partner can change. Settle where each of those will live before the contract is signed, not at handover. Our guide to agencies and consultancies goes further into what you own when an AI supplier's contract ends.

Ask two more things of any partner: who will actually do the work, and whether you can see one AI feature they have running in production, with the evaluation set behind it.

A partner alongside your own team

If AI is going to be part of what customers pay for, this is the route I would pick, with the same caveat as before: 1AYM is one of those partners. The partner builds the first feature with your engineers in the same repository and sets up the evaluation harness, the guardrails and the cost telemetry once. Your engineers build the second feature on those foundations while the partner reviews. By the third, the partner has stepped back and you know exactly which role to hire for.

It is slower than a partner working alone, because teaching takes time, and it only works if your engineers are given the hours. Put the handover in the plan with dates and names, or it becomes the thing that slips.

Evaluation is the part to own

An evaluation set is a collection of real inputs with agreed good outcomes, graded automatically, that tells you whether a change made the feature better or worse. Without one, every prompt edit and every model upgrade is a guess, and your customers do the testing.

The vendors' own guides agree on its shape. Anthropic's evaluation guide says to mirror the real mix of tasks, include edge cases and automate the grading, and it prefers many automatically graded cases to a few graded by hand. OpenAI's guide advises calibrating any model-based grader against human judgement and running the evaluations on every change.

Whoever builds it, the evaluation set should end up in your repository, run in your pipeline and be understood by your engineers. It is the asset that lets you change supplier, change model or change your mind without starting again.

What it costs to run, and who supports it

Model APIs are usually priced per token, so the running cost of a feature is roughly the tokens each request sends and receives, times the vendor's price, times how often customers use it. That number moves with every prompt change, every extra document retrieved and every customer who finds a heavier use than you planned for. OWASP's 2025 list of risks for LLM applications includes unbounded consumption for this reason. Measure the cost of each request from the first release, per customer, and decide in advance what happens when one customer's usage spikes.

Support is the cost that rarely makes the plan. When a customer disputes an answer, someone has to see which model version, prompt and retrieved documents produced it. When the vendor retires a model, someone has to run the evaluation set against its replacement. Those jobs fall to whoever owns the feature after launch, so decide who that is before you choose a route.

Your customers' data, and what you tell them

Before customer data goes to a model, read two documents you already have: your customer contracts and your privacy policy. In February 2024 staff at the Federal Trade Commission warned that it may be unfair or deceptive for a company to start using consumers' data for AI training and to tell them only through a surreptitious, retroactive change to its terms of service or privacy policy. The law behind that warning is Section 5 of the FTC Act, 15 U.S.C. § 45(a)(1), which declares unfair or deceptive acts or practices in or affecting commerce unlawful. If your customers are businesses, their contracts usually set out what you may do with their data, so read those first.

The same law covers what you say about the feature. Announcing Operation AI Comply in September 2024, the FTC's chair at the time said there is no AI exemption from the laws on the books. Describe what the feature does only as far as your evaluation results support. Where the model runs, and what the vendor may keep, depends on the hosting route; our guide to private LLMs, listed at the end of this page, compares the options.

This summarises US statute and FTC publications as they stood on 29 September 2026. This is not legal advice: have your own counsel review your contracts and your privacy policy.

How a partner should hand over

Handover is easy to promise and hard to do at the end, so it belongs in the contract as a set of named deliverables. These are the ones to ask for.

Your accounts from the first day
Repositories, cloud projects and model API keys sit in your organisation's accounts, with the partner added as users you can remove. Keys are rotated when the engagement ends.
The evaluation set and its harness
In your repository and running in your pipeline, with a short note on how to add a case when a customer reports a bad answer.
A runbook
What to do about a provider outage, a cost spike, a model retirement and a disputed answer, and whose job each one is.
A decision record
Which model was chosen and why, what was tried and rejected, and the evaluation results behind each choice, each entry dated.
A feature your engineers ship
Before the partner leaves, your engineers build a feature on the new foundations with the partner reviewing. A handover made only of documents tends not to stick.
Copyright in writing
A signed assignment of copyright in the code and documentation to your company, rather than reliance on a work-made-for-hire clause that the statute's categories may not support for software. Have your own counsel draft it; this is not legal advice.

Which route fits your situation

Start from the feature and the customers waiting for it, not from the staffing model you would prefer.

Use your own team when
The first feature is narrow, a person reads the output before acting on it, and someone senior has the time to own evaluation.
Hire when
AI will be a lasting part of what customers pay for, you know what the role will do in its first year, and someone in the company can judge the candidates.
Bring in a partner when
Customers need the feature before a hiring round could finish, the scope can be written down, and you have agreed where the code, the keys and the evaluation set will live.
Combine the two when
AI will be central to the product, you want your own team running it within the year, and you can give your engineers the hours to build alongside the partner.

If you cannot yet say which feature to build first, or for which customers, none of the four fits yet. Settle that before the staffing, because a hiring plan or a statement of work written before it is a guess.

Where 1AYM fits

1AYM builds AI features into production software and hands them over. On one government-accredited EdTech's live platform, about 1,170 automated tests cover the AI speaking assessment, among them a regression suite in which switching off any one guardrail breaks specific frozen cases. That is the kind of evaluation I would expect any partner to leave in your repository.

For a software company the usual shapes are a fixed-scope architecture and production build of the first feature, which can start within a day of the scope being signed, or engineers embedded in your team on a contract from three months, directed by you day to day and reviewed weekly by a 1AYM architect. If you already have the work scoped, 1AYM can resource it on contract from the associates who work with it, held to the same standard. The work is delivered in your repositories, copyright is assigned to you where your contract needs it, and US companies can contract through 1AYM's US entity.

For engineers: the build checklist for a first AI feature

Whichever route you choose, these are the pieces a first AI feature needs before it reaches customers. They are also the things to check in a partner's proposal, and the things that should be in your repository when the partner leaves.

An evaluation harness in CI
A versioned set of real inputs with expected outcomes, including the edge cases Anthropic's evaluation guide lists: irrelevant or missing input, overly long input, harmful input, and cases where people would disagree. Grade in code wherever the answer is exact, use a model-based grader only where it is not, and calibrate that grader against human labels, as OpenAI's guide advises. Run it on every change to a prompt, a model, a retrieval index or a tool.
Model versions pinned
Call a dated model version where the vendor offers one, rather than an alias that moves underneath you. Treat an upgrade as a code change: it ships when the evaluation set passes, and not before.
Untrusted input, untrusted output
OWASP's 2025 list for LLM applications puts prompt injection first (LLM01) and names improper output handling (LLM05) and excessive agency (LLM06). Validate structured output against a schema, escape it before rendering, and give any tool the model can call no more access than the user it is acting for.
Tenant isolation in retrieval
In a multi-tenant product, filter retrieval by tenant before the model sees anything. Sensitive information disclosure (LLM02) and vector and embedding weaknesses (LLM08) are both on the OWASP list. Keep a test case that tries to pull another tenant's documents, and fail the build if it succeeds.
Cost telemetry per tenant
Log input and output tokens per request, per feature and per tenant from the first release. Set per-tenant rate limits and a budget alert. Unbounded consumption (LLM10) is the failure mode this prevents, and it shows up on an invoice before it shows up anywhere else.
Fallbacks
Decide what the product does on a timeout, a rate limit, a provider outage and output that fails validation: retry, route to another model, or fall back to the non-AI path the feature replaced. Tell the user which one happened.
Logs that support can read
For each request, record the model version, the prompt template version, the IDs of any retrieved documents and the output, kept only as long as your privacy policy allows. That record is how support answers a customer who asks why the feature said what it said.

Sources

  1. [1]OpenAI, Evaluation best practices (API documentation)
  2. [2]Anthropic, Define success criteria and build evaluations (Claude documentation)
  3. [3]NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (July 2024)
  4. [4]OWASP GenAI Security Project, Top 10 for LLM Applications 2025
  5. [5]FTC Office of Technology, AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive (13 February 2024)
  6. [6]FTC, FTC Announces Crackdown on Deceptive AI Claims and Schemes (25 September 2024)
  7. [7]15 U.S.C. § 45, Federal Trade Commission Act, Section 5 (govinfo, 2023 edition)
  8. [8]US Copyright Office, Circular 30: Works Made for Hire (revised August 2024)
  9. [9]17 U.S.C. § 204, Execution of transfers of copyright ownership (govinfo, 2023 edition)

Frequently asked questions

Should a software company hire for its AI features or use a development partner?

Hire when AI will be a lasting part of what customers pay for and someone in-house can judge the candidates. Use a development partner when customers need the feature before a hiring round could finish and the scope can be written down. The two combine well: a partner ships the first feature alongside your engineers, and you then hire into a role you understand.

What is AI MVP development?

An AI MVP is the first version of an AI feature, built to test with real users before committing to a full build. The useful kind includes a small evaluation set from the start, even a few dozen real cases, because without one you cannot tell whether the second version is better than the first.

How long does it take to add an AI feature to a software product?

The model call is quick. The evaluation set, the guardrails, the cost controls and the support tooling set the timeline, and how much of each you need depends on the damage a wrong answer can do. A drafting aid a person reads before sending needs far less than a feature that changes records or advises customers. Ask for a plan that names all four, not only the feature.

Who owns the code when a partner builds our AI feature?

Under US copyright law, work an employee creates as part of their job is owned by the employer, but commissioned work is made for hire only in nine listed categories with a signed written agreement, and software is not named among them. So ask for a written assignment of copyright signed by the partner, which 17 U.S.C. § 204(a) requires for a transfer to be valid. This is not legal advice.

How do we estimate what an AI feature costs to run for each customer?

Multiply the tokens each request sends and receives by the vendor's per-token price, then by how often a typical customer uses the feature. Measure it per customer from the first release, because prompt changes, extra retrieved documents and heavy users all move it. Set a per-customer rate limit and a budget alert before launch, and ask any partner for this estimate at your expected usage.

Do we need our own AI team eventually?

If AI becomes part of what customers pay for, yes. Someone in-house has to own the evaluation set, the running cost and the model decisions. A partner engagement should end with your team able to do that, and the handover belongs in the contract as named deliverables.

Further

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM