Data volumes are growing exponentially. New data types, formats, and repositories are being created all the time. Data is being shared in the cloud and around the globe in millions of digital channels. Every electronic device—at home, at work, on the go—is constantly transmitting personal information, and companies need to locate and properly label and treat all of it.
Download this free ebook to learn more about:
- How to use data discovery and classification to gain greater visibility and understanding of your organization’s sensitive data
- Why thorough discovery is necessary to establishing a competitive advantage
- The simplest solution for discovering data wherever it lives and moves
Why Data Discovery Became a Requirement
Data volumes grow, and so do the number of formats and repositories holding them. Personal information moves through cloud services, shared drives, endpoints and applications that did not exist when the last inventory was taken.
Data management, information governance and regulatory compliance are the three drivers that push organizations toward automated discovery. In heavily regulated industries it is already treated as a necessity rather than an improvement.
Structured, Semi-Structured and Unstructured
Discovery is straightforward in a well-labeled database and difficult everywhere else. Personal data also lives in documents, exports, logs, spreadsheets and message archives, where no schema describes what a field contains.
That is why coverage is the measure that matters. A tool that only reads structured sources reports a clean result while the majority of the estate goes unexamined.
What the CPRA Added
California voters passed the California Privacy Rights Act in November 2020, extending the CCPA. It grants the right to limit the use, sharing and disclosure of sensitive personal information, the right to correct personal information, the right to know about automated decision-making and to opt out of it, and audit obligations around sensitive data.
It carries specific provisions for children’s data and for employee data, and it established the California Privacy Protection Agency, funded in part by the fines it collects.
What the Market Is Signalling
MarketsandMarkets put the global personal data discovery market at $5.1 billion in 2020 and projected $12.4 billion by 2026, a compound annual growth rate of 16.1 percent. Remote work and the demand for real-time data to feed machine learning are both cited as accelerants.
What to Look for in a Discovery Solution
Four things separate discovery tools in practice: the range of data types recognized, the techniques used to identify them, the quality of tagging and labeling produced, and how much configuration is required before the first useful scan.
Ready-made policies matter more than they appear to. A platform that ships with a library of them can be scanning on day one, whereas one that expects every rule to be written first delays the only output that matters, which is knowing what you hold.
Discovery Is the Step Everything Else Depends On
Classification, masking, encryption, retention, deletion and subject access requests all take the same input: an accurate account of what personal data exists and where it sits. None of them can be done well against a partial inventory.
That is why discovery is deployed first, and why its coverage sets the ceiling for everything built on top of it. A masking program applied to the data somebody remembered to scan protects exactly that data and nothing else.
Indexing Identities, Not Only Locations
Finding personal data in a repository is useful. Knowing which person it belongs to is what makes a subject access request answerable, and it is a different capability. An index linking identities to the records about them, across every repository, turns a right-to-know request from a search project into a query.

