I spend most of my week on calls with security teams at banks, insurers, and payers. Over the last year one question has changed shape. It used to be some version of “how do we classify data in Microsoft 365.” Now it’s “we turned on Copilot, and I’m not sure what it can reach.”
Those are different problems, and the second one doesn’t stay inside Microsoft.
What a sensitivity label actually is
Worth being precise here, because the term gets used loosely.
A sensitivity label is an instruction to a person, enforced by the applications that person works in. It says this document is restricted, handle it a certain way, and Purview makes that stick across Office, SharePoint, Teams, and Exchange. That’s a real control and it works.
It also assumes the reader is a human who had to go find the file. For fifteen years that was a safe assumption. The work happened in Office, the files sat in SharePoint, and getting to a document you weren’t supposed to see took effort and left a trail.
Anyone who built their program on that assumption made the right call. I’m not writing this to tell you the last decade was wrong.
What changed
Copilot honors permissions. That isn’t the failure, and I want to be clear about it because people expect a vendor to say otherwise. The failure is that those permissions were set by people who assumed retrieval was slow. Now it’s instant, and nobody has to know what they’re looking for.
That’s before you count everything that isn’t Copilot. Other assistants, third-party connectors, and sync tools all pull copies of data across the line Purview was drawn around. Machine identities now outnumber human ones by roughly 109 to 1, and around 40% of them can reach company data (Palo Alto Networks, 2026).
AI didn’t create this problem. It made the problem easy to find. A file with the wrong permissions in 2019 was a risk almost nobody tripped over. The same file today is one prompt away from someone who was never supposed to see it, and you hear about it after.
The estate is bigger than the admin console
Most of what follows never shows up in Purview reporting.
The regulated data I run into on customer calls is in a Box folder that came over with an acquisition. It’s in an S3 bucket a developer stood up for a project that ended. It’s on a laptop in a claims department. It’s in a database nobody has opened since the migration, and it’s in a z/OS dataset that has held customer records since before Teams existed.
Purview’s reach outside Microsoft is partial, and it takes real work to stand up. More to the point, its enforcement doesn’t follow the data once the data moves. It’s Microsoft’s control for Microsoft’s estate and it’s good at that job. Most of the rest stays unlabeled. Unlabeled isn’t the same as unimportant. It just means nobody has looked.
Recent Gartner® report mentions, “Furthermore, they should discover and classify sensitive data, leverage sensitivity labeling and establish Microsoft Purview as the foundational control — not necessarily the end state.” I read that gap as the distance between a control plane and an actual estate.
“AI didn’t create this problem. It made the problem easy to find.”
Where we fit in
Two layers, one policy.
Purview keeps its job. It decides what sensitive means inside Microsoft and it enforces there. We’re not asking anyone to rip it out. If a vendor is telling you to replace a control you’ve already deployed and trained people on, they’re solving their problem, not yours.
PK Protect covers the rest of the house.
We find sensitive data on endpoints, in cloud, in M365, on z/OS, and on IBM i. That last pair is the one I get asked about most, and it’s usually because the systems holding the oldest and most regulated records are the ones no tool ever scanned.
The same business definitions on every platform. Not one definition of sensitive for Microsoft, another for cloud, and whatever the mainframe team wrote in 2011. When an examiner asks how your company defines a Social Security number, there’s one answer, and it holds whether the record is sitting on a laptop or in a z/OS dataset.
When we find something, we can act on it. Encrypt, mask, redact, quarantine, delete. A discovery report with no remediation attached is a dated record showing when you found out.
Protection stays with the file after it moves. It doesn’t stop at the edge of the platform it started in, which is the only version of this that survives a copy or a sync nobody told you about.
Format-preserving encryption keeps the data usable. The application downstream still reads the field. That’s usually what decides whether an encryption project finishes or stalls in week three.
Discovery reports metadata, never file contents.
Now the part I’d want to hear if I were on your side of the call. We don’t monitor AI tools and we can’t tell you which agent touched which file. If that’s what you’re shopping for, we’re not it. What we do is make sure whatever those tools reach has already been found, classified, and protected. That’s the step before AI, and it’s the one most teams are skipping.
What I’d do first
Scan where visibility is worst, not where it’s easiest. You already have a rough idea what’s in SharePoint and OneDrive. You have close to none on endpoints, non-Microsoft cloud, and the mainframe. Easy scans produce clean reports about data you were already watching.
Settle the policy before you pick engines. Agree the business definitions with the people who own the data, then decide which tool enforces them where. The other order gets you four tools and four definitions.
Attach remediation to every finding. If discovery ends in a spreadsheet, all you’ve built is a record of the date you learned.
PKWARE has spent four decades on the unglamorous half of this, the finding and the protecting, on platforms that were unfashionable long before anyone called them critical. The teams that come through the next eighteen months clean won’t be the ones with the best AI policy. They’ll be the ones who knew where their sensitive data was before they turned anything on.
If you want to know what’s actually out there, we’ll show you. Request an assessment.
Gartner, How to Build a Data Security Program Using Microsoft Purview, Andrew Bales, 3 September 2026.
GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.
