Buyer guide
Data strategy for AI: what to fix before you build
A data strategy for AI is the short list of data decisions to settle before an AI system is built. For one use case at a time, it answers six questions: can the AI see only what the person using it may see, is the data accurate and current enough for this job, who owns each source and each business definition, is the data in documents or databases (which changes the preparation), is a document collection ready to be searched, and what UK data protection law asks if personal data is involved. It ends in a written decision for that use case: go, fix first or stop.
Data strategy for AI. The decisions about access, quality, ownership, format and personal data that make an organisation's data usable by a specific AI system.
Checked . The public guidance this page cites was read on each publisher's own site on this date. Guidance changes, and the ICO says its AI and DPIA guidance is under review following the Data (Use and Access) Act.
Start from one use case, not the whole estate
A data strategy for AI does not need to be a two-year data programme. It needs to answer one question well: can this data support this use case, safely? Answer it for the first use case, fix what that turns up, and the second use case inherits most of the work. Trying to make the whole estate AI-ready before anything is built is, in my view, how data programmes stall with nothing in production to show for them.
The UK government's AI Playbook puts the dependency plainly. AI systems rely heavily on the quality and quantity of their data, and poor quality, biased or noisy data can lead to inaccurate outcomes if it is not managed. The Playbook also warns that AI can make existing risks worse, and names poor data management and insufficient security classification among them.
One of our published engagements shows where the time really goes. Connecting a client's AI workspace to its finance data took a single afternoon. Agreeing the definitions that make an answer correct took a week with the finance director, and that week was the real work. Had we stopped after the afternoon, the client would have had a tool that looked finished and quietly disagreed with finance.
The six areas below are what to check, and the checklist at the end puts them on one page you can take into a meeting.
Access: the AI sees only what the person using it may see
Access comes first because it is the check that turns into an incident when it fails. The AI Playbook states the rule in one line: make sure your AI model only has access to the data the user of the model should be able to access.
The usual way to break it is a single service account that can read everything, with the AI trusted to work out what each person should see. A model is not an access control. The safer pattern carries each person's own permissions through to the data, so the AI cannot return anything the person could not have opened themselves. In the finance engagement, the connector carried each person's existing Looker entitlements, and nobody gained access they did not already have.
Respecting permissions exposes a second problem, which is oversharing. A file shared with the whole company years ago was hard to find by browsing. An assistant that searches everything a person can open will find it on the first day. Review sharing on the document stores in scope before you connect anything, rather than after the first surprised email.
Access also has to change when people do. Our engagement file on identity provisioning for more than 1,000 users shows the pattern: access follows the HR record, so a leaver loses it because a record changed rather than because somebody noticed. Test it with two people who have different access asking the same question. Each should get an answer built only from what they are allowed to see.
Quality: judged against the job the AI will do
The UK government's Data Quality Framework treats data quality as fitness for purpose: is this data set good enough for what I want to use it for? It is the right question for AI, because the bar moves with the job: a drafting assistant can live with gaps that would be unacceptable in a system that approves payments. The framework's six dimensions make a practical review when you take them one use case at a time.
| Dimension | What the framework means by it | What to check before an AI build |
|---|---|---|
| Accuracy | The degree to which data matches reality. | Sample the records the AI will rely on and check them against the source. For figures, compare them with the numbers finance or operations already report. |
| Completeness | The degree to which records are present. | Look for missing records and empty fields the use case depends on. An AI answering from half the contracts will answer confidently from half the contracts. |
| Uniqueness | The degree to which there is no duplication in records. | Find duplicate customers, suppliers and document copies. A duplicate in a document store is an old version the AI may quote. |
| Consistency | The degree to which values do not contradict other values representing the same entity. | Check that the same customer, product or figure agrees across the systems the AI will read. |
| Timeliness | The degree to which data reflects the period it represents and is up to date. | Know how often each source refreshes, and have the AI say how current an answer is where that matters. |
| Validity | The degree to which data is in the range and format expected. | Check formats, units and ranges, such as dates, currencies and codes, before they reach a prompt or a query. |
Fix problems at source where you can. The framework asks organisations to solve data quality issues at source rather than apply temporary fixes. A clean-up script placed in front of the AI is exactly that kind of temporary fix, and the next use case will have to write its own.
Ownership: a named person for every source and every definition
Every source the AI reads needs an owner who can say who may see it, what it means and when it is wrong. Without one, nobody can sign off the access rules above, and nobody can settle an argument about whether an AI answer is right.
Definitions are where ownership earns its keep. A business with three working definitions of revenue will get three kinds of answer from an AI, depending on which table it happened to read. Agree each definition once, write it down where the AI can use it, and name the person who is allowed to change it. In the finance engagement that person was the finance director, and the AI's answers were checked against his until the two agreed.
The same goes for documents. A policy with no owner and no review date is a policy nobody will notice is out of date until the AI quotes it to a customer.
Documents and databases need different work
AI use cases usually read one of two kinds of data, and the preparation differs. Documents, such as policies, contracts, case notes and wiki pages, are usually read through retrieval: the system searches for the passages that match a question and passes them to the model with it. The AI Playbook calls this retrieval augmented generation (RAG), and notes that it preserves access controls when the search covers only data the user is permitted to see. Databases, such as a warehouse, a finance system or a CRM, are usually read by the AI running a query, so the risk moves from finding the right passage to computing the right number.
| Dimension | Documents | Databases |
|---|---|---|
| How the AI reads it | Searches for matching passages and answers from them. | Runs a query, ideally against agreed metrics rather than raw tables. |
| What usually goes wrong | It quotes an old or draft copy, misses the passage that matters, or surfaces a file the person should not see. | It finds the right table and still returns the wrong number, because it chose a different definition or join. |
| What to fix first | One current version of each document, an owner and a review date, clear headings, and access tags carried with every passage. | Agreed definitions in a semantic layer or governed views, read-only access under each person's permissions, and checks on shape and freshness. |
| How to test it | Questions whose answers, and the passages they come from, the document owner has confirmed. | Figures reconciled against the system of record for periods where the right answer is already known. |
GOV.UK Chat, described in the AI Playbook, is a useful public example on the documents side. The team chose retrieval over fine-tuning because GOV.UK content changes regularly. Returning whole pages sometimes exceeded the model's token limit, so they worked on chunking, embedding models and re-ranking, and at an answer accuracy of 80% they concluded that accuracy had to improve before the product went to users. Most of the fixes they list are about how the content was split, indexed and ranked, which is where I would expect most of the effort on a document project to go.
On the database side, reconciliation is the same idea as our pattern for checking an AI system's output before it acts: the model proposes a figure or an action, and a deterministic check against the system of record decides whether it stands.
Retrieval readiness: preparing documents before an AI searches them
If the use case reads documents, run these checks on the collection before it is indexed. They are cheaper before indexing than after, because every fix afterwards means indexing again.
- One current version
- Archive or exclude superseded copies and drafts. Duplicates are the quickest route to a confident answer from last year's policy.
- An owner and a review date
- Each document has someone who confirms it is still right, and a date by which they will look at it again.
- Structure the search can use
- Clear headings and sections, so passages can be cut along the document's own structure rather than at arbitrary lengths.
- Access tags on every passage
- OWASP's guidance on retrieval systems recommends permission-aware vector and embedding stores, and tagging data in the knowledge base to control access levels. Filter by who is asking before anything is ranked.
- Trusted sources only
- OWASP recommends accepting data only from trusted and verified sources. A shared inbox or a folder that outsiders can write to is not one.
- Retrieved text treated as untrusted
- The AI Playbook warns that retrieval tools are susceptible to indirect prompt injection through the content they retrieve, so a document can carry instructions. Retrieved text should never trigger an action on its own.
- Logs of what was retrieved
- OWASP recommends detailed, immutable logs of retrieval activity, and the NCSC asks that logs are treated as sensitive data.
- A test set
- Real questions with answers the document owners have confirmed. The GOV.UK Chat team saw a knowledge base of quality-assessed questions as the route to semi-automated quality assurance.
Personal data: what UK GDPR asks before you build
If the data identifies people, data protection law applies to the AI system as it does to any other processing. The ICO's guidance on AI and data protection is the place to start, and the ICO says that guidance is under review because of changes made by the Data (Use and Access) Act. This section describes the guidance as published on 29 September 2026. It is not legal advice.
- Purpose and reuse
- Data collected for one purpose can be reused for a new one only if the new purpose is compatible with the original, and every new purpose needs a lawful basis. The ICO updated its purpose limitation guidance on 23 March 2026 for the Data (Use and Access) Act's provisions on compatibility and reuse.
- Building and using are separate
- The ICO says it will often make sense to separate the research and development phase of an AI system from its deployment when you settle purposes and lawful bases.
- Special category and criminal offence data
- Using AI on either brings the requirements of Articles 9 and 10 of the UK GDPR and the Data Protection Act 2018.
- Minimisation
- Personal data must be adequate, relevant and limited to what is necessary. The ICO is clear that this does not mean processing no personal data, and it describes ways to reduce what an AI system needs, including synthetic data and processing on the user's own device.
- A DPIA screen
- A DPIA is required where processing is likely to result in a high risk to people. The ICO lists innovative technologies, including AI, among the criteria that point to one when combined with others, so screen every AI use case that touches personal data.
- Where the data goes
- The AI Playbook's questions for any AI service: where your data is sent and how it is processed, whether it is used to train future models, how long it is retained, and who can read the logs. Our private LLM guide compares what each deployment option keeps private.
An AI-ready data checklist you can copy
Fill this in for one use case with the business owner and the data owners in the room. Every line you cannot complete is either a fix to make before the build or a reason to pick a different first use case.
AI-ready data checklist: [use case] Date: [date] Business owner: [name, role] Data owners: [name and role for each source] 1. Scope. The task or decision the AI will support: [one sentence]. The data it needs: [sources]. Data it does not need and will not be given: [sources]. 2. Access. The AI reads each source with the permissions of the person using it, not through an account that sees everything: yes / no. Sharing on each document store reviewed and overshared items fixed on [date]. Two people with different access tested, and each saw only their own data: yes / no. 3. Quality. Checked against this use case for each source: accuracy [ ], completeness [ ], uniqueness [ ], consistency [ ], timeliness [ ], validity [ ]. Known problems, and who fixes them at source: [list]. 4. Ownership and definitions. Every source has a named owner: yes / no. Business terms the AI must use (for example revenue, active customer, period) written down once and agreed by [name]. Where the definitions live: [location]. 5. Documents. One current version of each document, with older copies archived or excluded: yes / no. Every document has an owner and a review date: yes / no. Documents tagged with their access level or classification: yes / no. Only trusted sources indexed: yes / no. 6. Databases. The AI queries agreed metrics or governed views, not raw tables: yes / no. Answers reconcile with the system of record on [number] test figures: [result]. A change to an upstream schema or feed fails a check rather than changing answers silently: yes / no. 7. Personal data. Personal data involved: yes / no. Purpose written down, and any reuse checked as compatible with the original purpose: [reference]. Lawful basis: [basis]. Special category or criminal offence data: yes / no. Personal data the use case does not need removed or reduced: yes / no. DPIA screening done on [date]. Outcome: [not required / completed / in progress]. 8. Where the data goes. Where it is sent and processed: [location]. Used to train the supplier's models: yes / no. Retention: [period]. Who can read the logs: [roles]. 9. Test set. [number] real questions or cases with answers the data owners have confirmed. Rerun on every change to the data, the prompts or the model: yes / no. 10. Decision. Go / fix first / stop. What has to be fixed first, by whom and by when: [list].
The last line is the point of the exercise. A written fix-first list with names and dates beats a build that starts on data nobody has checked. For a broader read that also covers ownership, value and controls, the AI readiness assessment linked below takes a few minutes.
Where 1AYM fits
When the checklist shows the gaps are in the data layer, that is what our Data platform & AI enablement work covers: connecting AI to the warehouse under the permissions people already have, agreeing definitions once in a semantic layer, and adding data contracts and reconciliation so an AI answer lands on the same figures as the board pack. The finance engagement above is one example of it.
Where the use case is still open, a fixed-scope AI strategy sprint with a data and systems feasibility check ranks the candidates against your real data, usually over two to four weeks. Either can run as a statement of work, small or large, and start within a day of the scope being signed. A data job you have already scoped can be resourced on contract from the collective of associates who work with 1AYM, held to the same standard. Book a call below or email tayyeb@1aym.com.
For engineers: access, retrieval pipelines, queries and tests
The six areas as a technical reviewer would check them. Each item should be visible in configuration, code or logs, so none of it rests on anyone's word.
- Identity
- Delegated, per-user access to each source rather than a shared service account. Where a service identity cannot be avoided, filter results by the requesting user's entitlements at query time and log both identities.
- Permission-aware retrieval
- Store each document's access list or classification as metadata on every chunk and filter on it before ranking, in line with OWASP LLM08:2025. Re-sync permissions when they change at source, and test with two users of different access.
- Document pipeline
- Deduplicate, index the current version only, chunk on the document's own headings, and keep the source link, version, owner and review date with each chunk. Remove personal data the use case does not need before indexing.
- Structured data
- Query through a semantic layer or governed views with a read-only role and row-level security. Put data contracts on shape, meaning and freshness, and run a reconciliation job against the system of record.
- Untrusted content
- Treat retrieved text as input from outside the trust boundary. No tool call or write is triggered by retrieved content without a deterministic check, because the AI Playbook flags indirect prompt injection through retrieved data.
- Evaluation
- A versioned set of real questions with confirmed answers: expected passages for documents, expected values for figures. Run it in CI on every change to the pipeline, prompts or model, and fail the build below the agreed threshold.
- Logging
- One record per request: user, query, documents or tables read and their versions, answer and model version. Keep it access-controlled, since the NCSC guidance treats logs as sensitive data.
- Documentation
- For each data source: provenance, scope and limitations, retention time, review frequency and known failure modes, drawn from the documentation items the NCSC's secure development guidance lists for AI systems.
Sources
- [1]UK government, Artificial Intelligence Playbook for the UK Government (10 February 2025), read 29 September 2026
- [2]UK government, The Government Data Quality Framework (3 December 2020), read 29 September 2026
- [3]ICO, Guidance on AI and data protection (under review following the Data (Use and Access) Act), read 29 September 2026
- [4]ICO, How do we ensure lawfulness in AI?, read 29 September 2026
- [5]ICO, How should we assess security and data minimisation in AI?, read 29 September 2026
- [6]ICO, Principle (b): Purpose limitation (updated 23 March 2026), read 29 September 2026
- [7]ICO, When do we need to do a DPIA?, read 29 September 2026
- [8]NCSC, Guidelines for secure AI system development: secure development, read 29 September 2026
- [9]OWASP GenAI Security Project, LLM08:2025 Vector and Embedding Weaknesses, read 29 September 2026
Frequently asked questions
What is a data strategy for AI?
It is the set of data decisions to settle before an AI system is built: who may see what, whether the data is good enough for the job, who owns each source and definition, how documents and databases will be prepared, and what data protection law requires. Done well, it is scoped to one use case at a time and ends in a written go, fix-first or stop decision.
Do we need a data warehouse or a data lake before we can use AI?
Not necessarily. The first use case needs its own data reachable under controlled access, good enough on the quality dimensions that matter to it, and owned. A warehouse and a semantic layer earn their place when the use case needs figures from several systems to agree with the numbers finance reports.
How do we know if our data is good enough for AI?
Judge it against the job, not in general. The UK Government Data Quality Framework treats quality as fitness for purpose and measures it on six dimensions: accuracy, completeness, uniqueness, consistency, timeliness and validity. Check the dimensions that matter to the use case, then test the AI on real cases with answers the data owners have confirmed.
Should we fine-tune a model on our documents or use retrieval?
For documents that change, retrieval is usually the better starting point. The GOV.UK Chat team chose it over fine-tuning because its content is updated regularly, and the AI Playbook notes that retrieval can preserve access controls when the search covers only data the user is permitted to see. The Playbook also warns that a generative model can be made to reveal the data it was trained on.
Do we need a DPIA to use AI on customer data?
Often, yes. A DPIA is required where processing is likely to result in a high risk to people, and the ICO lists innovative technologies, including AI, among the criteria that point to one when combined with others. Screen every use case that touches personal data. This is not legal advice: your data protection officer or legal adviser should make the call.
How long does it take to get data ready for AI?
It depends on the gaps, and the connection is rarely the slow part. In one published 1AYM engagement, connecting a client's finance data to its AI workspace took an afternoon, and agreeing the definitions with the finance director took a week. Access reviews, ownership and data protection screening are usually what sets the pace, so start them first.
Further
- Data platform & AI enablement · Warehouse integrations, semantic layers, data contracts and reconciliation, so AI answers match the official numbers.
- AI readiness assessment · 13 questions on ownership, data, systems and controls, scored on the page with no sign-up.
- Engagement file D-06 · A finance data connector that kept every person's existing permissions, and the week that made its numbers right.
- AI strategy and roadmap · Ranking use cases against your real data and systems before anything is built.
- AI implementation: the first 30 days · What the first month should produce once the data is ready, week by week.
We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.
Last reviewed · 1AYM