The data AI needs most is the data we protect most closely

The evidence is growing that privacy-enhancing technologies can help us train cutting edge AIs on sensitive data – financial, healthcare, personal – while remaining ironclad in upholding privacy. Image: Getty Images
- AI needs recent, accurate and representative data, but that data is also the most sensitive and tightly regulated.
- Synthetic data can help with training, but real data is still required to build accurate models and make individual decisions.
- Privacy-Enhancing Technologies (PETs) let organizations learn from real data without exposing it or exchanging it.
As AI moves into finance, healthcare and fraud prevention, the debate is not about whether AI has enough data, but whether it has the right data.
Good models help lenders assess people with no credit history, banks detect fraud earlier and healthcare providers identify risk sooner. But good AI models are only as good as the data they learn from. The best data is recent, accurate and representative. The challenge is that it’s often also the most sensitive.
AI needs richer, more recent data
Useful data is often fragmented, sensitive and scattered across organizations. A bank may understand repayment behaviour, a retailer household spending, a telco access to digital services. Each sees only a fragment of any individual's life.
The value for AI lies in how those signals relate to one another. Combined safely, they reveal financial resilience, vulnerability, fraud risk or an unmet need more accurately than any dataset alone.
Recency matters too. Behaviour, economic conditions and fraud tactics change constantly, and models trained on outdated information are making decisions based on a world that has already changed.
The data AI needs most is the hardest to reach
Combining those datasets into something usable is where the difficulty begins. Financial data can reveal income, debt and vulnerability. Healthcare data can reveal medical history. Retail and telco data reveal everyday behavior and preferences.
That sensitivity is why real data cannot be shared freely, even when the purpose is valuable. Most data supply chains still rely on copying and transferring data across third-party systems, where a breach may occur or data collected for one purpose may later be used for another. The overriding issue is that once shared, data cannot easily be unshared.
Regulation imposes limits. Purpose limitation and data minimization mean data collected for one purpose generally cannot be repurposed for another, however valuable. The obstacle is not carelessness: responsible organizations conclude, correctly, that the risk of moving sensitive data between them outweighs the benefit.
How the Forum helps leaders make sense of AI and collaborate on responsible innovation
Can synthetic data solve the problem?
Synthetic data - or artificially created data designed to reflect the patterns of real-world datasets - looks like a logical way out. If the data is not real, it follows that the privacy problem should disappear. The truth is that synthetic data has a genuine role: testing systems, exploring scenarios and stress-testing rare or extreme cases, but it cannot carry the full weight of high-stakes AI.
Synthetic data only reproduces what is already known, and may miss emerging behaviors, unexpected correlations or subtle characteristics that exist in live data but not in the assumptions used to generate it. It cannot create what was never there. Synthetic data is generated from real data, so where real data is thin, the synthetic version is thin in the same places and will not look thin. Gaps and bias are inherited.
And training is not the same as deciding. Even a model trained entirely on synthetic data must, at the point of decision, assess an applicant, a transaction, a patient.
Germany's Federal Ministry of Finance has reportedly proposed allowing tax authorities to use real, unaltered taxpayer data to develop AI systems, arguing that fictitious data was ineffective; critics raised concerns about privacy and function creep.
Whether to use synthetic data or real data comes down to the use case. But for high accuracy of the decisions that matter most, real data is critical.
PETs offer a safer way to use real data
Privacy-enhancing technologies (PETs) are a set of technologies that facilitate data analysis without needing to exchange or expose sensitive data.
Records are matched using tokenized identifiers, so two organizations can establish which customers they share without either learning the identity of the other’s customers. Analysis then takes place inside a neutral environment where raw records are never visible to the other party or the operator. Only results leave the environment, such as aggregate insights, model features or a score for a single decision.
PETs do not remove the need for consent, accountability, or regulation. But they address the central obstacle to collaborating on real-world data.
What this looks like in practice
PETs are being used right now to solve the problem of sensitive data analysis. In South Africa, the continent's largest grocery retailer and its leading banks tested whether grocery shopping behavior could predict creditworthiness, without consumer data being shared between them.
Anonymized and tokenized datasets were matched inside a privacy-preserving environment and models trained on the combined data. Eight million previously credit invisible people were scored, 3.2 million qualified for affordable credit they would otherwise have been declined and models using grocery data showed lift in accuracy of 41%.
This was only possible because finance businesses could access anonymized grocery data via PETs. None of the evidence used existed in any bank’s records, and no synthetic process could have produced it: the behavior had never been observed by the institutions making the decisions. The data existed. It was just out of reach because of justified privacy concerns.
The AI-data obstacle is no longer technical
Adoption of PETs has been slowed by cost, complexity and the absence of common standards. Those barriers are receding as the techniques mature and regulatory guidance grows clearer.
PETs do more than make existing data safer to access and collaborate on. They expand what data can be used at all in exactly the areas where models are weakest: thin-file consumers, emerging fraud patterns, underserved patient populations.
Better AI does not depend only on better algorithms or larger models. It depends on access to real, recent, representative data about people without asking them to surrender their privacy. Privacy-enhancing technologies are how both become true at once.
Don't miss any update on this topic
Create a free account and access your personalized content collection with our latest publications and analyses.
License and Republishing
World Economic Forum articles may be republished in accordance with the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Public License, and in accordance with our Terms of Use.
The views expressed in this article are those of the author alone and not the World Economic Forum.
Stay up to date:
Data Policy
Forum Stories newsletter
Bringing you weekly curated insights and analysis on the global issues that matter.
More on Artificial IntelligenceSee all
Charles Pacini, Irene Varoli and Nicola Mackow
August 31, 2026



