Buyer guide
Machine learning or generative AI: which does your problem need?
Generative AI is a branch of machine learning, so the practical choice is between two kinds of system. A predictive model is trained on your own past records to return a number, a score or a category: a demand forecast, a fraud score, a churn risk. A generative model is a large model someone else trained in advance, which reads and writes language, images or code: drafting, summarising, answering questions over documents, pulling fields out of messy text. Choose predictive machine learning when you hold labelled history and there is one right answer to measure against. Choose generative AI when the input or output is language and you have nothing to train on. Many useful systems use both, with a rule or a person making the final decision.
Machine learning vs generative AI. Machine learning is the branch of AI that learns patterns from data to predict or classify; generative AI is the part of it whose large pre-trained models produce new text, images or code.
Checked . The government guidance, standards and role definitions this page cites were read on their publishers' own sites on 29 September 2026. The examples are illustrations, not client results or benchmarks, and no vendor or price is compared. This page is not legal advice.
Generative AI is machine learning, so where is the line?
The question usually arrives as a choice between two technologies, and that framing is slightly off. The UK government's AI Playbook defines machine learning as the branch of AI that learns from data, and says plainly that modern large language models are also examples of machine learning systems. Generative AI, in the Playbook's words, is a subset of AI capable of generating text, images, video or other forms of output by using probabilistic models.
The line that matters to a buyer sits between two kinds of system you could build. The first is a predictive model trained on your own records to answer one narrow question, such as which invoices are likely to be paid late. The second is a large model that someone else trained in advance, which you call through an API and point at your documents and workflows. For the rest of this page I call the first predictive machine learning and the second generative AI, because that is how suppliers sell them.
They differ on almost everything a sponsor cares about: the data you need before you start, where the money goes, how the system fails and how you prove it works. They also need different people. That is why it matters whether a supplier's recent work trained models or called them.
Start from the output you need
The quickest test is the shape of the answer. If the output is a number, a probability or a label from a fixed list, and you hold past cases where you know how things turned out, you are looking at predictive machine learning. If the input or the output is free text, a document or an image, and there is no single correct answer to train on, you are looking at generative AI.
| Dimension | Usually fits | Why |
|---|---|---|
| Forecast a number: demand, cash, staffing | Predictive machine learning, or plain statistics | You have the history, and every forecast can be checked against what happened. You need a number you can backtest, and a language model is the wrong shape for that. |
| Score a risk: fraud, churn, late payment, lead quality | Predictive machine learning | The answer is a probability you can set a threshold on, test against past outcomes and explain to an auditor. |
| Spot the unusual: transactions, sensor readings, system logs | Predictive machine learning (anomaly detection) | It runs cheaply on every record and learns what normal looks like from your own data. |
| Route or tag short text at volume: tickets, emails, feedback | Either; often generative AI first, predictive later | A language model can sort text into categories with no training data, which gets you started. Once you hold thousands of labelled examples and the volume is high, a trained classifier is usually cheaper per item. |
| Pull fields out of documents: invoices, contracts, forms | Generative AI, with validation | Layouts vary and the text is unstructured. Check every extracted field against rules or a system of record before anything is written. |
| Draft, summarise or answer questions over your documents | Generative AI | There is no single right answer to train on, and the value is in handling language. Ground it in your own sources and test it on real questions. |
| Recognise things in images: defects, damage, counts | Computer vision, a kind of predictive machine learning | Trained on labelled images of your own products or sites. The Playbook notes that older computer vision systems still thrive because they are simple, stable, well proven and light on computing power and memory. |
| Decide something about a person: eligibility, hiring, credit | Neither on its own | A model can inform the decision, but rules and an accountable person should make it. The Playbook puts fully automated decision making under its use cases to avoid, and asks for caution wherever a decision is significant. |
| A task a written rule already handles | Neither | The Playbook asks teams to stay open to the conclusion that AI is not the best solution, because a problem may be more easily solved with more established technologies. |
Decisions about people also raise data protection questions, about fairness and about how accurate the system is, which the ICO's guidance on AI and data protection covers. The ICO says that guidance is under review following the Data (Use and Access) Act, so check the current version with your data protection officer before the design is fixed. This page is not legal advice.
When the answer is both
Many of the systems worth building combine the two, and the combination is often safer than either on its own. The usual split gives the language model the unstructured part of the job and hands the decision to something that can be tested case by case.
Take supplier invoices that arrive as PDFs. A language model reads each one and pulls out the supplier, the amounts and the purchase order number, and rules then match those fields against the purchase order. A predictive model trained on past invoices flags the ones that look unusual, and a person approves those. Each step has its own test, and at no point does the language model have to be right about money on its own.
That shape, where the model proposes and a deterministic check decides, is the one I reach for whenever a wrong answer costs money or affects a person. It also works the other way round: a language model can turn free text into inputs, such as the topic of a complaint, that a predictive model then uses alongside your structured data.
What each one needs from your data
Predictive machine learning needs history before it can do anything. You need past cases with the outcome recorded, enough of them to cover the situations the model will meet, and a definition of the outcome that did not change halfway through the data. The Playbook's warning applies directly: AI systems rely heavily on the quality and quantity of data, and insufficient data can leave a model failing to generalise. If the outcome was never recorded, the first project is collecting it.
Generative AI needs no training data to start, which is why a prototype can appear within days and why demos flatter it. It still needs three things from you: the documents or records it will read, access controls so it sees only what the person using it may see, and a set of real cases with known good answers to test it against. I would rarely start with fine-tuning a model on your own data. Putting the right documents in front of it usually matters more.
Either way, check access, quality, ownership and personal data for the one use case before anything is built. A feasibility answer that arrives after the build is just an expensive one.
What drives the cost and the risk
Neither approach is cheaper in general. They spend money in different places and fail in different ways, so a business case written for one will miss the costs of the other. The table compares the drivers rather than prices, because prices depend on your volume, your data and the scope.
| Dimension | Predictive machine learning | Generative AI |
|---|---|---|
| Where the upfront money goes | Finding, cleaning and labelling historical data, and preparing the inputs the model learns from. | Integration with your systems, getting the right documents to the model, and building a test set of real cases. |
| What running it costs | Usually little per prediction once trained. Retraining and monitoring are the recurring lines. | Usually charged per use, so it grows with volume and with the length of what the model reads and writes. |
| How it fails | Drift: the world changes, the patterns in old data stop holding and the model gets quietly worse. | Confabulation, which NIST defines as confidently stated but erroneous or false content, delivered in the same tone as a correct answer. |
| The security risk particular to it | Training data that is wrong, biased or tampered with, which the model then learns as if it were true. | Prompt injection: instructions hidden in a document or message that change what the model does. OWASP puts it first in its 2025 list of risks for LLM applications. |
| Explaining a single answer | Simpler models can be explained input by input. The Playbook notes that with deep learning it can be hard to trace how an input led to an output. | The answer can read like reasoning. The Playbook warns that LLMs are not domain experts and are no substitute for professional advice in legal, medical or other critical areas. |
| What you depend on | A model you trained, which you can run for as long as it stays accurate. | Called through an API, a provider's model versions, which change and are retired on the provider's timetable. A self-hosted open-weights model moves that upkeep to your own team. |
Both columns share the costs that business cases tend to leave out: an owner, monitoring, and a re-test every time the model, the data or the prompt changes. Put those in before comparing the two, or the cheaper-looking option will simply be the one with fewer lines counted.
How you prove each one works
Predictive models have the easier evaluation problem, because there is a right answer. Hold back data the model never saw in training, ideally the most recent months, and measure it there. Compare it with the simple method it would replace, such as last year's figure or the current rule, because a model that cannot beat that is not worth running.
Pick the measure that matches the cost of being wrong. The ICO's guidance on AI separates precision, the share of cases flagged as positive that really are positive, from recall, the share of real positives that get flagged, and notes the trade-off: chasing recall can cost you false positives. A fraud model that misses little but flags a great deal will bury the team that checks its output.
Generative AI is harder to test, because there is often no single right answer. Build a set of real cases from your own work, a few hundred where you can, write down what a good answer must contain, and grade the output against it, with people checking a sample. Include cases designed to go wrong, such as documents carrying hidden instructions and questions your sources cannot answer.
Do not buy on public benchmark scores. NIST's profile for generative AI warns that measurement gaps arise from mismatches between laboratory and real-world settings, and that tests restricted to benchmark datasets may not carry over to real conditions. A model's score on a public test tells you very little about your invoices.
Both need checking after launch as well as before. The Playbook asks for full testing before deployment, regular checks of the live tool, and a way for users to report problems that prompts a human review.
Which specialist to hire
The two approaches call for different people, and the Playbook says so directly: developing bespoke AI solutions and training your own models require different specialist skills from using pre-trained models through APIs. Hiring the wrong profile is an expensive way to find that out.
| Dimension | Who you need | What to ask them |
|---|---|---|
| A predictive model trained on your records | A data scientist to establish whether the data supports it, then a machine learning engineer to put it into production and keep it working. | What simple method will the model have to beat? How will you test it on data it has not seen, and how will you notice drift? |
| A generative AI system on your documents or workflows | An engineer who builds on pre-trained models: integration, retrieval, access controls and a test harness. A data scientist is optional at the start. | Show me the test set and how it is graded. How do you stop prompt injection, and what will each case cost to run at our volume? |
| Both, or you are not yet sure | Someone who owns the architecture and makes the call before either specialist starts. | Which parts of this need a trained model, which need a language model, and which need neither? |
The government's Digital and Data Profession Capability Framework draws the same line. It describes a machine learning engineer as someone who develops, assures and maintains machine learning models so they can be used in products and services, and at senior level deploys them into production and checks that live models stay safe, secure and continue to work. Its data scientist explores data using statistical tools and techniques, such as machine learning and predictive analytics, to inform decisions.
The same goes for suppliers. A firm selling machine learning consultancy and a firm selling generative AI can each be strong on their half and thin on the other. Ask each for the last system it put into production of the kind you need, and how it proved that system worked. If the problem needs both, you want one party accountable for how the pieces fit, which is the job described in our guide to what an AI solution architect does.
Where 1AYM fits
For a business weighing several possible uses, 1AYM's AI Opportunity & Feasibility Sprint is where I would settle this. It is fixed scope and typically takes two to four weeks: we look at how the work runs today and what your data and systems can support, score the candidate use cases on value, feasibility and risk, settle build versus buy, and set out the technical architecture for the recommended direction. You come away with a plan an engineer can build from and a sponsor can fund.
Where the answer is a system built on pre-trained models, with the controls that let it touch real data and real decisions, we build that too, as fixed-scope production AI systems work. 1AYM takes small fixed-scope statements of work as well as larger builds, and if you already have a scoped job, we can resource it on contract from the collective of associates who work with us, held to the same standard. If you want a second opinion on which kind of problem you have, book a call below.
For engineers: baselines, evaluation and hybrid patterns
These are the choices where projects in both camps usually go wrong. None of them depends on a particular vendor or model.
- Baseline first
- For tabular prediction, fit a regularised logistic regression or a gradient-boosted tree model before anything deeper, and record its score. For a generative feature, record what the current manual or rules-based process achieves on the same test cases. Every later model has to beat those numbers.
- Splits that respect time
- Split by date, not at random, when the data has a time order, and hunt for leakage: any input that would not have been known at the moment of prediction. Leakage produces the best offline scores and the worst live ones.
- Calibrated scores
- If a person or a rule acts on a threshold, check calibration as well as ranking, so that a score of 0.8 means roughly eight in ten. Set the threshold from the cost of a false positive against a false negative, not from a default.
- Structured output from language models
- When a language model extracts data, ask for output against a fixed schema, validate every field in code, and reject anything that fails or route it to a person. Never pass free text from the model straight into a write to a system of record.
- Evaluation harness in CI
- Keep the graded test set in the repository with its thresholds, and rerun it on every change to a prompt, model version, retrieval setting or training set. Store the outputs, so a regression can be traced to the change that caused it.
- Embeddings as features
- A cheap hybrid: turn free text into embeddings or model-extracted labels, then train a conventional classifier or regressor on them alongside the structured columns. Inference is fast, and the result can be tested like any other predictive model.
- Monitoring
- For predictive models, compare live input and score distributions with training, and track the realised outcome once it is known. For generative systems, sample live outputs for human grading, and log every retrieved document and tool call so a bad answer can be reconstructed.
- Injection controls
- Treat every document, email and web page the model reads as untrusted input. Give the model's tools least-privilege access, keep writes behind a deterministic check, and put injection cases in the test set. OWASP notes that retrieval and fine-tuning do not fully mitigate prompt injection.
Sources
- [1]UK government, Artificial Intelligence Playbook for the UK Government (10 February 2025), read 29 September 2026
- [2]NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (July 2024), read 29 September 2026
- [3]OWASP, LLM01:2025 Prompt Injection, OWASP Top 10 for LLM Applications 2025, read 29 September 2026
- [4]ICO, Guidance on AI and data protection: What do we need to know about accuracy and statistical accuracy? (updated 15 March 2023; under review following the Data (Use and Access) Act), read 29 September 2026
- [5]Government Digital and Data Profession Capability Framework, Machine learning engineer (last updated 28 August 2026), read 29 September 2026
- [6]Government Digital and Data Profession Capability Framework, Data scientist (last updated 29 August 2025), read 29 September 2026
Frequently asked questions
Is generative AI a type of machine learning?
Yes. The UK government's AI Playbook defines machine learning as the branch of AI that learns from data, and names modern large language models as examples of machine learning systems. In practice, buyers and suppliers use "machine learning" for predictive models trained on an organisation's own data, and "generative AI" for large pre-trained models that produce text, images or code.
Can a large language model make predictions from our data?
It can be asked to, and it will return a number. For a forecast or a risk score on structured records, a model trained on your own history is usually cheaper to run, easier to test against what actually happened and easier to explain. Use the language model for the parts that are language, such as reading documents or turning notes into structured fields.
Do we need a data scientist for a generative AI project?
Not usually at the start. The first need is an engineer who can integrate a pre-trained model, control what it can see and do, and build a test set to measure it. A data scientist earns a place once there is data to analyse, such as graded outputs or usage at volume, or if the project turns out to need a trained model as well.
What is the difference between a machine learning consultancy and a generative AI consultant?
A machine learning consultancy typically builds predictive models from your data: forecasting, scoring, anomaly detection and computer vision. A generative AI consultant typically builds on large pre-trained models: document processing, assistants, drafting and search. Some firms do both. Ask any supplier for the last production system of the kind you need, and how it was tested.
Which is cheaper, machine learning or generative AI?
Neither in general. Predictive machine learning puts more of the cost up front, in preparing and labelling data, and is usually cheap to run. Generative AI is quicker to start because it needs no training data, but more of its cost comes from use, so it grows with volume. Compare the two at your own volume and over the life of the system.
Further
- AI strategy and roadmap · The fixed-scope sprint that scores use cases on value, feasibility and risk and settles build versus buy.
- AI solution architect · The role that owns the call when a problem needs a trained model, a language model or both.
- Data strategy for AI · What to check on access, quality, ownership and personal data for one use case, before either kind of build.
- AI ROI: estimate it, then check it · Where the running, monitoring and re-testing costs belong in the business case for either approach.
- Hiring an AI engineer · Employ, contract, bring in a team or buy a build, once you know which specialist you need.
- Production AI systems · The fixed-scope architecture and build for a system on pre-trained models.
We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.
Last reviewed · 1AYM