SynthProof

Independent synthetic data validation

Prove your synthetic data is faithful, private, and useful.

Upload a dataset. Get a scored certificate covering fidelity, privacy attacks, and downstream utility — benchmarked against your category. We are the auditor, not the generator.

One email when we launch. No newsletter, no sharing your address.

Validation certificate
Adult census · SDV Gaussian Copula
75/100
Grade B
Fidelity
Statistical similarity to the source data
85/100B
Privacy
Resistance to re-identification attacks
98/100A
Utility
Train on synthetic, test on real
35/100D
Structural integrity
Internal consistency of the synthetic data
99/100A
Metrics engine v0.4.0 · source data deletedRead the full report

A real run, published unedited — including the pillar it failed.

One of these datasets is a copy of a real production table.

Both were scored by the same engine. The copy wins on fidelity, wins on utility, and looks structurally clean — every quality metric a generator computes about its own output would have called it excellent.

SDV generatorMemorising generatorUCI Adult census · 4,200 rows × 15 columns
Pillar scores out of 100 for an SDV generator and a generator that memorised its training data, on the UCI Adult census dataset.
PillarScore out of 100SDVMemorised
Fidelity
A copy matches the source perfectly, by definition.
84.8100.0
Utility
Models trained on real records transfer to real records.
35.169.3
Structural integrity
Both datasets are structurally clean.
98.897.4
Privacy
4,200 of 4,200 synthetic rows were byte-for-byte copies of a real record.
98.68.0
Overall grade75.4 B54.0 D
▲ marks where memorising scores better. A generator that returned its training data verbatim beat a real one on fidelity and utility, and looked structurally clean. Every metric a generator computes about itself would have called it excellent. The privacy pillar found that 4,200 of its 4,200 rows were byte-for-byte copies of a real person's record, and the critical-finding rule capped its grade at D.

Generators grade their own homework

Every synthetic data tool ships with its own quality metrics. Those metrics are computed by the same system that produced the data, and they are the numbers you are asked to show your DPO, your auditor, or your client.

We are not in the generation business and never will be. That is the whole point: the certificate is worth something precisely because we have nothing to gain from your dataset scoring well.

Four questions, answered with evidence

Each pillar is measured with published, reproducible methods — SDMetrics for distributions, Anonymeter for the Article 29 attack criteria, scikit-learn for downstream performance.

Fidelity

Do the distributions and the relationships between columns survive generation? Per-column shape tests and every pairwise correlation, not just the flattering ones.

Privacy

Can anyone get a real person back out? Exact-copy detection, nearest-neighbour distance against a real baseline, membership inference, and singling-out, linkability and attribute-inference attacks.

Utility

Does a model trained on the synthetic data still work on real data? We train on synthetic, test on real, and compare against the same model trained on real data — measured above chance, not from zero.

Benchmarking

How does this compare to everyone else in your category? Percentiles against a frozen benchmark snapshot, with the snapshot version stamped on the report.

Every report says what to look at, not just what you scored

A score tells you a pillar is weak. A finding tells you which column, what happened to it, and the kind of fix it calls for — so the report is something your team can act on rather than a number to argue about.

Named columns, not a verdict

Findings name the column whose distribution drifted, the pair whose relationship was lost — or invented, which is worse — and the column driving your privacy risk. Where a relationship changed, the report prints the correlation on both sides rather than a similarity score you have to decode.

Diagnosis, never a sales pitch for a fix

We say what is wrong and the general direction of a remedy. We never tell you which knob to turn in the tool that produced the data, and we never offer to fix it for you. An auditor that also sells you the remedy is not an auditor — that boundary is the product.

A clean dataset gets none

Findings are derived from the measurements, not from your data, so they survive the deletion stamped in the report footer. And they only appear when there is something to say: a report with no findings is a result, not an empty section.

From the sample report

1 column distribution diverges from the source

capital-gain

Weakest is capital-gain at 9/100 distribution match; 1 column scores below 70.

What to look at: Per-column drift usually means the generator under-fitted these marginals — common with heavy tails, rare categories, and columns whose type was inferred wrongly upstream. Confirm each column is being treated as the type it actually is, and check whether rare categories survive generation at all.

One of 3 findings on that run. Read the whole report.
Distributions
SDMetrics
Privacy attacks
Anonymeter
Attack criteria
Article 29 WP
Utility
TSTR · scikit-learn
Every report stamped
engine + benchmark version
Every input
SHA-256 fingerprinted
Read the full method — every metric, threshold and guarantee

How it works

You sample your own data before upload. No database connectors, no VPC peering, nothing of ours touching your systems.

Mode 1 · full report

Synthetic data + a real sample

Upload your synthetic dataset alongside a representative sample of the source — 10k to 100k rows, never the full production table. All four pillars run.

Optionally add a holdout: real records your generator never saw. That is what turns the attack metrics from upper bounds into measurements, and it is the only way membership inference can be tested at all.

Mode 2 · partial report

Synthetic data only

Can't share real data at all? We score structural integrity — duplicates, collapsed columns, impossible values, missing-data rates — and issue a partial report.

Fidelity, privacy and utility are comparisons against real data, so they are marked not assessed with the reason. They are never shown as a zero, and never as a pass.

Zero retention, every plan

Your source data is deleted by the worker as the final step of every job — success or failure — and the deletion timestamp is printed in the report footer. This is the default on the free tier and every paid tier alike, not an enterprise add-on.

File bytes go straight from your browser to storage via a presigned URL; they never pass through our application servers. There is an opt-in checkbox to retain data for 24 hours if you want us to help debug a failed job. It is off by default.

Because the data is gone, the report carries the SHA-256 of every file we measured. That fingerprint is what ties a set of scores to a specific file, permanently.

Pricing

Usage-based, no seats, no sales call for the self-serve tiers. Processing region — US or EU — is a choice on every paid plan, not something bundled into the top tier; the free plan runs in the US.

Starter
Free
  • 1 dataset / month, up to 50 MB
  • Fidelity + utility scoring
  • Basic privacy leakage checks
  • Watermarked certificate
Pro
$299/month
  • 20 datasets / month, up to 2 GB
  • Full privacy attack suite
  • Percentile benchmarking
  • Clean certificate + API access
Compliance
€799/month
  • 50 datasets / month, up to 10 GB
  • DPIA-mapped report language
  • Signed processing statement / DPA
  • Everything in Pro

Talk to us — design partners first

Get the first certificate

We're opening validation to a first group of teams. Leave your email and we'll get in touch when your category is ready.

One email when we launch. No newsletter, no sharing your address.

Or read a full sample report first — it's a real run, published unedited.