MOSTLY AI vs. Tonic vs. Syntheticore: which one fits your data?
All three show up on the same shortlists, and they're not really competing for the same job. MOSTLY AI (now relaunched as "MOSTLY AI, powered by Syntho") and Syntheticore both generate privacy-safe tabular data with a published fidelity-and-privacy report; Tonic.ai is built for engineering teams who need de-identified, referentially intact test data and PII-safe text. The scorecard below compares all three on deployment, data types, privacy metrics, pricing, and open source.
Facts about vendors below were verified on October 2, 2026. This market moves fast — check current documentation before you commit.
One of these three already had its "Gretel moment"
If synthetic-data vendors feel like they're consolidating out from under you mid-evaluation, you're not imagining it. MOSTLY AI — the Vienna company, founded in 2017, behind one of the category's most-cited tabular synthesizers — shut down in March 2026 after raising more than €30 million. Two months later, on June 9, 2026, Syntho (a Netherlands-based synthetic-data vendor) announced it had acquired the MOSTLY AI brand and trademark, and the product now continues as "MOSTLY AI, powered by Syntho." The open-source Synthetic Data SDK (Apache 2.0) the original team released in February 2025 still works and is still on GitHub; what the public announcement doesn't say is what the acquisition means for existing hosted-platform customers, support, or roadmap under the new owner.
We covered the same pattern — a category name outliving the company behind it — in our guide to Gretel's NVIDIA acquisition. The lesson applies here too: ask any vendor, including us, what a wind-down or acquisition would mean for your data, your contracts, and your support.
What each one actually does
MOSTLY AI, powered by Syntho
Best for: teams that want a privacy-safe tabular synthesizer with a genuinely
open-source core — the Synthetic Data SDK is Apache-2.0 licensed and runs locally via pip,
with an enterprise Data Intelligence Platform on top for teams that want a hosted or self-managed UI.
Watch for: pricing and roadmap are now Syntho's to set following the brand acquisition,
and the public self-serve tiers that existed before the March 2026 shutdown are gone from the current
site — plan on a sales conversation, and ask directly about support continuity if you were a pre-acquisition
customer.
Tonic.ai
Best for: engineering teams that need de-identified, referentially intact copies of
production databases for dev, test, and QA (Structural); PII redaction across free text, documents, and
images in 50+ languages (Textual); or schema-first synthetic data generated with no source data at all
(Fabricate). Of the three, it's the only one with a public, self-serve, free-to-start tier.
Watch for: its headline privacy claims are benchmarked differently from the other two —
PrivacyBench measures how accurately PII is detected and replaced (precision, recall, F1), not
distance-to-closest-record or membership inference. If your reviewer wants the latter, ask explicitly; it's
different evidence for a different kind of risk.
Syntheticore
This is us, so weigh it accordingly. We generate privacy-safe tabular synthetic data — including
multi-table datasets with keys and joins intact — through a web app and a REST API, deployable as managed
SaaS or inside your own cloud. Every run ships a utility and privacy report (train-synthetic/test-real,
distance-to-closest-record, membership inference), and we pair the platform with people who help land the
first dataset.
Watch for: no public self-serve pricing, no free tier, and no free-text or image
generation — if that's the job, look at the other two.
The scorecard
Feature grids age badly and every vendor can tick every box in a demo — score these on your data using the seven questions in our buyer's guide. For a starting point, here's what's publicly verifiable today:
| Criterion | MOSTLY AI powered by Syntho |
Tonic.ai | Syntheticore |
|---|---|---|---|
| Deployment | Open-source SDK, local; Data Intelligence Platform on Kubernetes/OpenShift, self-hosted or cloud | Self-hosted (VPC, on-prem, air-gapped) or Tonic-hosted cloud, per product | Managed SaaS or inside your own cloud, via web app and REST API |
| Data types | Tabular — single- and multi-table/relational with keys, plus sequential data. No free text or images | Relational/NoSQL databases and files; free text, documents, images in 50+ languages; greenfield generation with no source data | Tabular, including multi-table datasets with keys and joins. No free-text generation |
| Privacy metrics | Auto-generated report per generator: distance-to-closest-record, DCR share, nearest-neighbor distance ratio, identical-match share, discriminator AUC | PrivacyBench — an open benchmark scoring PII detection and replacement in text (precision/recall/F1), not a distance or inference measure | Report on every run: train-synthetic/test-real utility, distance-to-closest-record, membership inference |
| Pricing model | No public self-serve tier since the Syntho relaunch; contact sales. SDK itself is free | Mixed: Fabricate is self-serve (free, then $29/mo, custom Enterprise); Structural is custom-quoted; Textual is usage-based plus Enterprise | Contact sales; no public self-serve tier |
| Open-source options | Yes — Synthetic Data SDK (Apache 2.0) plus a benchmarking toolkit | Partial — several components on GitHub (Textual's engine, PrivacyBench metrics, a RAG-eval tool), not the full platform | No open-source release today |
Which one actually fits you
- Provisioning dev/test/QA environments from production databases → Tonic Structural.
- Redacting PII from documents, tickets, or chat logs for an LLM or RAG pipeline → Tonic Textual.
- Generating synthetic data from a schema with no source data yet → Tonic Fabricate.
- Running a tabular synthesizer yourself, for free, with the code in hand → MOSTLY AI's open-source SDK — with eyes open about the current pricing and roadmap uncertainty around the hosted platform.
- A managed platform with a published utility-and-privacy report on every run, plus hands-on help landing the first dataset → Syntheticore.
Frequently asked questions
Is MOSTLY AI still an independent company?
No. The original Vienna-based MOSTLY AI shut down in March 2026 after raising more than €30 million. Syntho acquired the MOSTLY AI brand and trademark in June 2026 and continues the product as "MOSTLY AI, powered by Syntho." The Apache-2.0 open-source SDK the original team released still works; the hosted platform's pricing, support, and roadmap are now Syntho's to set.
Does Tonic.ai do the same kind of synthetic data as MOSTLY AI or Syntheticore?
Not quite. Tonic's center of gravity is test data — de-identified, referentially intact database copies (Structural) and PII redaction in free text (Textual) — plus schema-first generation with no source data (Fabricate). MOSTLY AI and Syntheticore are built around privacy-safe synthesis of an existing tabular dataset for analytics and ML. There's overlap, but each is optimized for a different job.
Which of the three has public pricing?
Only Tonic Fabricate publishes self-serve tiers (free, then $29/month, as of its mid-2026 pricing page). Tonic Structural and Textual, MOSTLY AI post-acquisition, and Syntheticore are all sales-led with no public price list.
How do their privacy metrics actually differ?
MOSTLY AI and Syntheticore report distance-based and membership-inference-style metrics on generated rows — evidence that no synthetic record sits suspiciously close to a real one. Tonic's PrivacyBench instead scores how accurately a model detects and replaces PII in text (precision, recall, F1). Both are legitimate evidence; they answer different questions, so match the metric to the risk you're actually managing.
Which one open-sources the most?
MOSTLY AI's Synthetic Data SDK is fully open source under Apache 2.0, including its benchmarking toolkit. Tonic has open-sourced individual components — the Textual detection engine and the PrivacyBench metrics among them — without open-sourcing the full platform. Syntheticore doesn't currently publish open-source code.
The bottom line
These three aren't really three answers to the same question. If your job is engineering test data or scrubbing PII from documents, Tonic's lineup is the most complete. If you want a tabular synthesizer you can run yourself under a permissive license, MOSTLY AI's SDK still works, though the business around it is mid-transition. If you want a managed platform that treats privacy as a number on every run and backs it with people, that's the case we'd make for Syntheticore. Whichever you pick, measure utility and privacy on your own data before you commit — see how to measure synthetic data — and ask every vendor the acquisition question before you need the answer.