The essentials in 30 seconds
- Synthetic data are realistic decoys that trigger your controls just like real data would, without being real.
- They make testing possible in production: if the decoy gets out, it carries no real business data. The operational risks are managed through scope.
- A good decoy matches your formats, carries a traceable marker and is never derived from production.
“Too risky in production.” It is one of the objections we hear most often, and on the substance, the CISO is right. Nobody should accept pushing real customer data out just to check that it gets blocked. Our answer comes down to one idea: keep the test in production, and change what goes out. If the file that leaves is a decoy, no real business data goes out with it, and the operational risks that remain are managed through scope.
Synthetic data are realistic decoys (fake IBANs, fake customer records, fake documents marked “confidential”) designed to trigger your controls as real sensitive data would, without being real. They are what allows us to test in production, where the results are representative.
Why test in production, not just in pre-production
A DLP never behaves in acceptance testing the way it does in production. Your workstations do not run the same image, your proxies do not carry the same exceptions, your cloud policies have drifted since the March review, and your users have installed tools nobody declared. Testing elsewhere, in other words, means validating an environment the attacker will not target. Our offer follows the same two-step logic: a POC in a synthetic environment, with no impact on production, to see the tool, then a pilot in production on a restricted scope, with synthetic data only, to measure.
The page Data exfiltration testing: definition, method and stakes sets out the overall framework.
How synthetic data changes the risk
The risk of an exfiltration test comes down to 3 questions: what can get out, where it can land, and who must be told if it does. Decoys change the answer to each one.
| Risk | Test with real data | Test with synthetic data |
|---|---|---|
| Actual leak | Potential personal data breach, to be assessed | No real personal data, business data or secrets used as payload |
| Uncontrolled destination | Residual copy whose deletion cannot be guaranteed | Worthless copy, traceable by its marker |
| Notification obligations | Analysis with the DPO, notification to the supervisory authority (the ICO in the UK) within 72 hours if the breach poses a risk to individuals (Article 33 of the GDPR and the UK GDPR) | No real personal data involved |
| Internal acceptance | Hard trade-off, often refused | Discussion focused on workload and scope |
Designing good decoys
A decoy your DLP does not recognise tests nothing. A decoy that looks like a random string skews the result: if it gets through, you do not know whether the rule has a gap or simply ignored implausible data. Put simply, the quality of your test depends on the quality of your synthetic files.
Diagram · anatomy of a good decoy
Match the formats your rules detect
The decoy must reproduce the structures your rules look for:
- Identifiers with a check digit: a fake IBAN (ISO 13616 standard, 22 characters for a UK account, check digits calculated modulo 97) or a fake NHS number (10 digits, the last one a modulus 11 check digit) must have a valid key. Otherwise a well-written rule ignores it, and rightly so.
- Payment card numbers: use published test numbers, such as 4111 1111 1111 1111 (Visa) or 5555 5555 5555 4444 (Mastercard). They pass the Luhn algorithm without matching any real account.
- Business documents: the Python Faker library produces plausible names and addresses in British English, American English and many other locales, ready to drop into a fake Word contract, a fake PDF payslip or a fake patient record, with the layout and vocabulary of your real documents.
- Classification labels: if your rules rely on labels or metadata, the decoy must carry them. Without consistent data classification, some of your rules never fire, and the test should show it.
- Plausible volumes: a 10-line file and a 50,000-line export do not stress your thresholds in the same way.
Give every decoy a traceable marker
Each piece of synthetic data must carry a unique marker that answers one question unambiguously: “this file seen in a log, where does it come from?”. Here is what we recommend:
- a unique identifier per scenario and per run, such as ENF-TEST-2026-031, written into the content and the metadata;
- a central register where you link each marker to its date, the channel tested and the result;
- an explicit test notice, readable by an analyst, so that a decoy found later is never handled as a real incident.
With this marking, your SOC can measure whether the event was detected, how many minutes it took, and which tool caught it. It is the foundation of a robust time-to-detect metric.
Never derive from real data
A concrete scenario: an HR file to a personal cloud
In a hospital, on a Thursday, you want to check that a payroll file cannot leave an HR department workstation for a personal cloud storage service.
- You generate a 200-row synthetic Excel spreadsheet: invented names, fake National Insurance numbers in a valid format, fictitious salaries, a “HR Confidential” label.
- You record it in the register with its marker.
- We replay the scenario from a real Windows workstation in scope, under a standard user profile, to a consumer cloud storage service such as Google Drive. In MITRE ATT&CK terms, this is technique T1567.002.
- You record 3 results: was the transfer blocked, was an alert raised, did the SOC handle it (and in how many minutes)?
- If the file gets through, you know which control gave way. And the file that left contained no real business data.
Next step
Would your controls stop a synthetic payroll file?
A 30-minute demo: a full campaign across the 8 channels we test, on our demo environment. The POC shows, on site, how the solution behaves in your environment, on one simple scenario in a synthetic environment. What actually leaves through your channels is what the pilot shows you.
Answering the “too risky in production” objection
Faced with a reluctant DPO, you need a written framework. Here is what I would put in it:
- Nature of the data: generated synthetic data only, never derived from production.
- Destinations: listed in advance. With Enforcis 1.0, these are GitHub Gist, Google Drive and DPaste test accounts controlled by Enforcis and purged after each campaign.
- Scope: your named workstations, accounts, channels and time slots (Tuesday from 10 am to 12 noon, for example), approved by the IT department.
- Notification: who is told and when, for example 48 hours beforehand (the whole SOC and the on-call team, or just one point of contact if you want to measure detection).
- Stop: a reachable contact and a procedure to suspend a scenario in under 5 minutes.
- Traceability: a register of markers, kept and shared with the DPO.
With this written framework in place, the discussion with the DPO is about scope and schedule. No longer about the risk of a leak. And it is in production that you will find DLP false negatives: the rule exists, but it does not fire on the real workstation.
Limits to keep in mind
Synthetic data does not replace everything. It tests your controls’ ability to recognise known structures. It says nothing about sensitive data you have never described in a rule. And it ages: a new contract template missing from your decoys creates a blind spot. Keep your decoy library up to date with your document templates.
FAQ
Can synthetic data trigger a notification to the ICO?
No, provided it contains no real personal data and is not derived from production. Its exfiltration is not a personal data breach within the meaning of Article 4(12) of the GDPR or the UK GDPR, so there is no 72-hour notification. Still, document the approach with your DPO.
Can you test exfiltration in production without any risk?
Not without any risk, but without using any real business data: only synthetic payloads go out. Operational risks do remain. A poorly calibrated scenario can set off a burst of alerts or slow down a workstation. We frame them in writing: named workstations, accounts, channels and time slots (Tuesday from 10 am to 12 noon, for example), SOC informed 48 hours beforehand, and the ability to stop in under 5 minutes.
Are decoys enough to test the whole DLP?
They cover rules based on content, formats and labels. Combine them with a review of classification and of the channels covered to test your DLP thoroughly.
Can we use real card numbers, masked?
No. A masked number is still tied to a real account, and it often no longer passes your rules’ format checks. Use the published test numbers for Visa or Mastercard. That is what they are for.
