A translucent shield around a cloud of data points, with a few points at the edge outside it
Guide · 9 min read · Last updated Oct 10, 2026

Is synthetic data covered by GDPR, HIPAA and the CCPA?

Not automatically. Under GDPR, HIPAA and the CCPA, synthetic data is not a legal category of its own: it is treated as anonymous or de-identified only if the result can't reasonably be linked back to a person, and that is a question of evidence about your dataset and generator, not of the label. This guide covers what each regime says, where synthetic data can still count as personal data, and the evidence a data protection officer will ask for.

This is not legal advice. It is a technical team's reading of public rules, written to help you prepare for a conversation with your counsel and your DPO. Laws, regulator guidance and court decisions change and differ by jurisdiction; confirm anything you rely on with a qualified lawyer. Last reviewed October 10, 2026.

The short version

GDPR: anonymous information is out of scope, but the bar is high

GDPR applies to "personal data": any information relating to an identified or identifiable natural person (Article 4(1)). Recital 26 says the Regulation does not apply to anonymous information, and that to decide whether a person is identifiable you consider "all the means reasonably likely to be used", by the controller or by anyone else, taking into account cost, time and available technology.

The Article 29 Working Party's Opinion 05/2014 on anonymisation techniques is still the usual reference for what "anonymous" requires. It tests for three risks: singling out (isolating a record that belongs to one person), linkability (connecting records about the same person across datasets) and inference (deducing an attribute of a person with significant probability). Data that resists all three is anonymous; data that doesn't, including pseudonymised data (Article 4(5)), remains personal data.

GDPR doesn't mention synthetic data, so a generator's output is judged by that same test. Two consequences follow:

Regulators have generally treated well-built synthetic data as a promising privacy-enhancing technology while stopping short of calling it automatically anonymous. Read the current guidance from your own supervisory authority; several have published or updated anonymisation guidance in recent years.

HIPAA: two de-identification routes, neither names synthetic data

HIPAA's Privacy Rule applies to protected health information (PHI) held by covered entities and their business associates. Health information that has been de-identified under 45 CFR 164.514 is no longer PHI. The rule offers two methods:

Synthetic data usually fits the second route more naturally than the first. Safe Harbor is a checklist about the content of fields, and a synthetic dataset can still contain a rare combination of values that re-creates a real patient. Expert Determination is a documented risk analysis, which is where the metrics later in this guide do their work. Creating the synthetic dataset from PHI is itself a use of PHI that must be permitted under the Rule, and HHS's de-identification guidance is the place to confirm how the expert's role and documentation are expected to look.

CCPA/CPRA: "deidentified" comes with promises, not just math

California's Consumer Privacy Act, as amended by the CPRA, defines personal information broadly, as information that identifies, relates to, describes or could reasonably be linked with a particular consumer or household. It treats "deidentified" information, meaning information that cannot reasonably be used to infer information about, or otherwise be linked to, a particular consumer, differently, provided the business also:

So for the CCPA, evidence that your synthetic data is not linkable is necessary but not sufficient: you also need the public commitment and the contract terms. Separately, the statute carves out some data already governed by other regimes, such as HIPAA-covered PHI and certain medical information, so a healthcare dataset may be governed by HIPAA rather than the CCPA. Other US states have their own definitions; check each one that applies to you.

Where synthetic data still counts as personal data

Synthetic data is only as safe as the generator and the checks behind it. These are the common ways it stays in scope:

For a longer treatment of why masking alone fails the same test, see synthetic vs. anonymized data; for the broader privacy picture, see how synthetic data protects privacy.

The evidence a DPO will ask for

No law sets a numeric threshold for "anonymous enough". What a DPO or privacy lawyer needs is a defensible, documented assessment. These are the artifacts that tend to carry it.

Distance to closest record (DCR)

For each synthetic row, the distance to its nearest real training row, compared against the distance from a held-out real row to the training set. If synthetic rows are systematically closer to training data than genuinely new real rows are, the generator is copying. Report the distribution, the share of exact matches (which should be zero for most tables) and the share of rows closer than the holdout baseline.

Nearest-neighbor distance ratio (NNDR)

The distance to the closest real record divided by the distance to the second-closest. A ratio near 1 means a synthetic row isn't specifically tied to one individual; a very small ratio means it sits almost on top of one person's record. It catches the case DCR can miss, in a dense region where a row is close to many people but suspiciously tied to one.

Membership-inference testing

Train an attacker model to guess, from the synthetic data alone, whether a given real record was in the training set, and measure it on members against non-members held out from training. Accuracy or AUC near chance (0.5) is the goal; a materially higher score is evidence of leakage. Include the strength of the attacker you tested against, because "no leakage against a weak attack" proves little.

The rest of the file

Our own platform produces a utility and privacy report on every run, including distance-to-closest-record and membership-inference results, so there's something concrete to hand the reviewer. That report supports the legal assessment; it doesn't replace it.

A practical checklist

  1. Establish a lawful basis and purpose for using the source data to train the generator.
  2. Decide which regime or regimes apply to the data (GDPR, HIPAA, CCPA, others) and who your counsel is.
  3. Choose a generator and settings with privacy in mind, and consider differential privacy for the most sensitive data.
  4. Run DCR, NNDR and membership-inference tests against a holdout, and keep the outputs.
  5. Write down the threat model, the thresholds you accepted and why.
  6. Get the sign-off in the form your regime expects: DPO review, Expert Determination, or documented reasonable measures plus the CCPA commitments and contract terms.
  7. Re-run when the data, generator or audience changes.

Frequently asked questions

Is synthetic data exempt from GDPR?

Only if it is anonymous under Recital 26, meaning no one can reasonably identify, link or infer information about a person from it using the means reasonably likely to be used. Synthetic data that leaks training records or membership is still personal data, and creating it from personal data is processing that needs a lawful basis.

Is synthetic data HIPAA de-identified?

Not by default. HIPAA recognises Safe Harbor and Expert Determination (45 CFR 164.514(b)). A synthetic dataset derived from PHI is de-identified only if it meets one of them, and Expert Determination, a documented, expert-reviewed risk analysis, is usually the better fit.

Does the CCPA treat synthetic data as deidentified?

It can, if the data can't reasonably be linked to a consumer or household and the business also takes reasonable measures, publicly commits not to re-identify it and binds recipients by contract. The statistical evidence alone doesn't satisfy those conditions.

What privacy metrics should I show a DPO?

Distance to closest record, nearest-neighbor distance ratio, exact-match counts and membership-inference results against a held-out set, plus a written threat model and utility results. Add attribute-inference or linkage tests when the data holds sensitive fields or joinable quasi-identifiers.

Is there a "safe" DCR or NNDR threshold?

No law or regulator sets one. Thresholds are a judgment about risk for your data and audience, which is why the assessment should compare against a holdout baseline, state the threat model and explain the accepted values.

Do I still need a lawful basis to create synthetic data?

Yes, in the GDPR context. Training a generator processes the source personal data, so you need a lawful basis and a purpose-compatibility assessment where the data was collected for something else. Under HIPAA the creation must also be a permitted use of PHI.

Is this legal advice?

No. It is general information from a synthetic-data company. Have your own counsel and DPO review your use case against the law and guidance that apply to you.

The bottom line

Synthetic data can take a dataset out of scope of GDPR, HIPAA or the CCPA, but only when you can show it, and "we used a synthetic data tool" is not showing it. Treat the legal question and the measurement question as one project: pick the regime, run the distance and inference tests, write down the threat model, and get the sign-off your regime expects. To see how that looks in practice, see how synthetic data protects privacy and how to measure it.

Talk to us about privacy evidence for your data → More from the blog