Shadow AI is already using your data. Get the complimentary Gartner® report. Read the report

Data Privacy

Somebody is going to ask you to prove your masked data can’t be reversed

An auditor, a regulator, or a customer's third-party risk team is going to ask you to prove that a masked dataset can't be put back together.

PKWARE

By PKWAREProductivity Protected

Share on social media

Not describe the method. Prove it. In writing, with the testing attached, to someone who has read the current guidance and knows what re-identification research looks like in 2026.

If your honest reaction to that is “we’d have a rough two weeks,” you’re in the majority. Most data protection programs were built when masking a column was the end of the conversation. You ran the job, the field came back with X’s, somebody signed a memo, and the dataset went in the pile marked safe.

That model is running out of road. Two things arrived at the same time to push it there.

The legal definition is being tightened

On July 7, 2026, the European Data Protection Board adopted draft Guidelines 02/2026 on Anonymization and opened them for public consultation through October 30. They’re a draft, not binding, and the technical detail can still move. They replace guidance written in 2014, which is older than most of the attack research it was meant to survive.

Two things in the draft matter more than the rest.

01

Anonymity is treated as relative, not absolute.

The same dataset can be personal data in one organization’s hands and anonymous in another’s, depending on who holds it and what they could realistically combine it with. That follows the Court of Justice in EDPS v SRB last September. The practical effect is that you can’t certify a dataset as anonymous in the abstract. You certify it for a recipient, a context, and a point in time.

02

Documentation moves to the center.

The draft expects controllers to document how the data was anonymized, including the testing done on the supposedly anonymous output, and to keep that documentation afterward. Most of the law firms who’ve published readings of the draft land in the same place: a re-identification risk assessment becomes a standing artifact rather than a one-time memo.

Waiting for the final text is a legitimate call. It’s a much better call if you already know what sensitive data you’re holding and where it sits. Most teams don’t, and that’s the part that takes quarters, not weeks.

AI moved the same problem into unstructured data

Here’s what turns a European legal question into a Tuesday problem in Milwaukee.

The data your teams want to put in front of a model isn’t the tidy customer table. It’s the contract PDF, the claims file, the support ticket, the call transcript, the shared drive nobody has audited since the last migration. That’s where the sensitive material actually lives, and it’s exactly what a masking program built around database columns was never designed to touch.

“LLMs now drive demand for unstructured data masking and synthetic data.”

Source: Gartner®, Market Overview: Data Masking and Synthesis, Joerg Fritsch, Andrew Bales, Anson Chen, 19 August 2026.

Meanwhile the ask changed shape. It used to be “give the test environment a safe copy.” Now it’s “give the model everything, and give it now, because our competitors already have.” Legal says no. The AI team hears that as a failure of nerve. Neither side is being unreasonable, and nobody in the room has the evidence to end the argument.

That bind is the real problem. You’re being asked to enable something fast and prove something slow, and the proof layer was never built.

What it costs is specific. A discovery report with no remediation attached is a dated document proving you knew. A model pointed at an unreviewed folder becomes a finding you hear about from somebody else. And the question you can’t answer in the room, the one where somebody asks whether the data going into the AI program is clean, is the one people remember about you.

“You’re being asked to enable something fast and prove something slow, and the proof layer was never built.”

Where we fit

Plainly, so nobody has to guess: PK Protect finds sensitive data across endpoints, cloud storage, Microsoft 365, and the mainframe, classifies what it finds, and masks or encrypts it in place. Discovery with the remediation attached, on the surfaces most tools stop short of.

We’d rather be straight about this market than sell you a story about it. A masked field and a formal, mathematically stated privacy guarantee are two different things, and the industry is moving toward measurable re-identification risk, synthetic data, and formal privacy methods. Any vendor telling you all of that is solved and shipping everywhere is selling.

What we will say for our own part is that none of it works if you can’t find the sensitive data first. Finding it across a mainframe, a laptop, and a SharePoint site in the same program is the part most teams still can’t do, and it’s the step every other conversation depends on.

Start there. Know what you have, where it lives, and what’s inside it. Every conversation about masking, synthesis, or a privacy guarantee is one you can’t have honestly until that question is answered.

Start there. Know what you have, where it lives, and what’s inside it.

Book a data discovery assessment and see what PK Protect finds across your endpoints, cloud storage, Microsoft 365, and the mainframe.

PKWARE was named as an Example Vendor in Gartner®, Market Overview: Data Masking and Synthesis, Joerg Fritsch, Andrew Bales, Anson Chen, 19 August 2026.

Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.