← All articles

ARTICLE · DATA LEAK ACTIVE TESTING

Choosing a DLP testing tool: criteria and questions for the vendor

  • Buyer’s guide
  • 6 min read
  • Updated
Row of red warning triangles on a circuit board, an image of the criteria to check before choosingThree knock-out criteria

The key points in 30 seconds

  • A DLP testing tool measures your DLP by replaying exfiltration attempts; it does not replace it.
  • 3 criteria are knock-out criteria: synthetic data, testing in production, observable proof for each scenario.
  • Send our 16 questions in writing, then insist on a 4-week pilot that replays the scenarios after remediation.

“Their demo on Tuesday was impressive, the tool works.” We often hear this after a market review, and it mixes up two things. A demo shows the tool in an environment prepared by the vendor, with rules it knows by heart. It says nothing about your DLP, your exceptions or your estate. Ours lasts 30 minutes and is no exception: it shows you the product, not your leaks.

A DLP testing tool replays controlled exfiltration attempts to check that your data leak prevention rules block what they are supposed to block. When choosing one, one criterion comes before all others: its ability to produce verifiable evidence, scenario by scenario, without using any real business data. The rest (coverage, frequency, deliverable) follows from that.

What a DLP testing tool must prove, and what it must not promise

Put simply, a testing tool does not replace your DLP, it measures it. Its value lies in one question: “would this sensitive data, through this channel, from this workstation, have got out today?”. If the answer comes as an abstract score or as configuration compliance, you have not bought a test. You have bought a configuration audit.

The fundamentals are covered in our page on data exfiltration testing; here, we stay focused on buying.

The criteria grid

Diagram · three filters before the pilot

The three knock-out criteria before a pilotThree successive filters: harmless data, using synthetic data only; testing in production, on real workstations and the real network; observable proof, blocked, detected or exfiltrated. A tool that fails any of them is rejected. Those that pass all three go to pilot, then are compared on coverage, frequency and deliverable.1Harmless datasynthetic data onlyrejected2Testing in productionreal endpoints, real networkrejected3Observable proofblocked, detected or outrejectedPilot, then comparisoncoverage, frequency, deliverable

The table below ranks 8 criteria by weight. The first 3 are knock-out criteria: a tool that fails any of them does not deserve a pilot.
Criterion What to check Warning sign
Harmless data (knock-out) Exclusive use of synthetic data designed to trigger your rules The tool asks for a sample of real data “for calibration”
Testing in production (knock-out) Runs on your real workstations and your network Testing only in a lab or on a mock-up
Observable proof (knock-out) For each scenario: blocked, detected or exfiltrated, with a timestamp Results limited to an overall score
Channel coverage The channels you need, listed explicitly, with what is not covered Coverage presented as complete, without a channel list
Frequency On-demand replay, plus scheduled or change-triggered replay Annual campaign billed as a consulting engagement
Deliverable Prioritised corrective actions, each linked to a finding Bulky report with no prioritisation
Tracking over time History, regression detection, export to the SIEM Every test starts from scratch
Deployment effort Windows and Linux agents, required privileges, deployment options (vendor cloud, private cloud, on-premises) Extensive admin rights with no justification

Why harmless data comes before coverage

A test that handles real data creates the very risk it claims to measure: if the channel is not blocked, the data really does get out. Well-built decoys (correctly formatted IBANs, card numbers that pass the Luhn check, documents tagged with your labels) trigger the same rules with no consequences. See our article on synthetic data in production.

Why production matters

In practice, a DLP rule that is perfect in pre-production can fail in production for mundane reasons: an exclusion added in Purview for a business service, an agent missing from part of the estate, a proxy bypassed by a thick client. These gaps produce silent false negatives, and they only show up in the real environment.

The questions to ask the vendor

Send these 16 questions, grouped into 4 blocks, in writing before any demonstration. A written answer is more binding, and it lets you compare vendors line by line. Ask everyone, including us.

On proof

  1. For a given scenario, what does the result show: the channel, the synthetic data used, the control that reacted, the time?
  2. How do you distinguish “blocked”, “detected but exfiltrated” and “exfiltrated with no alert”?
  3. Can you correlate the result with the logs from our DLP or our SIEM, Splunk or Microsoft Sentinel for example?
  4. How do you handle a scenario with an ambiguous result?

On the security of the tool itself

  1. What data does the tool collect on our workstations, and where is it stored?
  2. What privileges do your agents require on Windows and Linux, and why?
  3. How do you prevent a third party from hijacking the tool for a real exfiltration?
  4. Who on your side can access our results?

On coverage

  1. Which channels do you cover today, including the Chrome, Edge and Firefox browsers, and which are on the roadmap?
  2. Are your scenarios mapped to MITRE ATT&CK, for example to the 9 techniques of the TA0010 tactic, from T1020 (automated transfer) to T1567 (web services), by way of T1048 (alternative protocols) and T1052 (physical media)?
  3. How do you incorporate new techniques observed among attackers?
  4. Do you cover everyday usage (personal Dropbox, Teams or Slack, Copilot or Gemini) or only advanced techniques?

On operations

  1. How long between installation and the first actionable result?
  2. Who prioritises the corrective actions, and on what criteria?
  3. How do you measure that a fix has worked?
  4. How do you present progress to non-technical management?

Specialist or general-purpose platform

Breach and attack simulation (BAS) platforms cover the whole intrusion chain, and exfiltration is just one module among many. A specialised tool covers fewer tactics, but goes further on egress channels, data formats and interaction with your DLP rules. If your question is about data leakage, specialisation is worth it. See our comparison of BAS and exfiltration testing.

The same logic applies to pentesting and red teaming. These are useful one-off exercises that give you a snapshot. A continuous testing tool gives you a film. See also our comparison of continuous validation, pentest or red team.

A four-week pilot scenario

Do not sign without a pilot. Here is a reasonable format for a mid-sized company or a local authority:

  1. Week 1, from Monday: deployment on a representative scope (one department, a few dozen workstations, one cloud tenant). Measure the real effort.
  2. Week 2: first full replay, with synthetic data only. Compare the results with what your team thought it was blocking.
  3. Week 3: apply 2 or 3 priority fixes suggested by the tool.
  4. Week 4: replay. Check that the fixes hold and that no regression has appeared, then review on Friday.

If at the end you cannot explain in 5 minutes what was getting out, what no longer gets out and what remains to be dealt with, the tool is not doing its job. That applies to us too.

Next step

What does proof of effectiveness look like, scenario by scenario?

A 30-minute demo to see the tool, then your POC to judge it on the evidence.

What to check before deciding

  • The tool never uses real data.
  • Every result is linked to a channel, a control and a timestamp.
  • The channels that matter to you are covered today.
  • Replay after remediation can be launched without the vendor’s involvement.
  • The deliverable prioritises actions instead of just listing them.
  • Your security team has approved the privileges the tool requires and the data it collects.

FAQ

Which criteria rule out a DLP testing tool?

In our view, the answer comes down to 3 criteria. The tool uses only synthetic data and never asks for a real sample “for calibration”. It runs in production, on your real workstations and network, not on a mock-up. Finally, it delivers timestamped proof for each scenario: blocked, detected or exfiltrated. A tool that fails any of the 3 does not deserve your 4-week pilot.

Can you test your DLP yourself, without a tool?

Yes, occasionally, with test files and a few channels. The limits appear quickly: maintaining scenarios, covering the whole estate and replaying regularly takes time that few teams have. The manual method is described in how to test your DLP and find out whether it really blocks.

How do you check the security of the testing tool itself?

Ask 4 questions in writing. What data the tool collects on your workstations, and where it stores it. What privileges its agents require on Windows and Linux. How the vendor prevents a third party from hijacking it for a real exfiltration. And who, on the vendor’s side, can access your results. Have your security team validate the answers, ours included. As for Enforcis: agents in your environment, orchestration in the Enforcis cloud or in your private cloud; 100% on-premises deployment possible, subject to assessment, for critical environments.

Do you need a POC before the production pilot?

It is useful if your team wants to see the tool without touching production. The Enforcis POC runs in a synthetic environment. The pilot follows, on a limited production scope, with synthetic data only, a report, prioritised findings and a retest.

Put this grid to us

With synthetic payloads, we run controlled campaigns on the configured paths and observe how your controls actually behave, scenario by scenario. Send us the grid above: we will answer it in writing.