Between websites, social media, communications, mobile, the Internet of Things, and more, there’s an astounding amount of data created every day. According to one study, humans created 2.5 quintillion bytes of data daily in 2020. That number will only keep growing: Some experts predict that the amount of data people create each day will reach 463 exabytes by 2025.
As businesses use their data, they also create metadata, which is data in the context of who, what, when, where, why, and how. This information helps users identify the content of the larger data, including what it might mean and how it can be used. Data and, by extension, metadata will continually be on the rise, challenging data governance processes to keep up so that data is both usable and protected.
Read this free ebook to learn more about:
- The importance of the data governance foundation to protecting metadata
- Why data security governance is key to successful data analysis using metadata
- How to protect and support metadata governance framework for full-breadth data security governance across the enterprise
What Metadata Actually Is
Metadata is data placed in context: who, what, when, where, why and how. It tells a user what a larger body of data contains, what it might mean and how it can legitimately be used. The shorthand is data about data, and it is usually produced by cataloging and classification tools.
Its value is practical. Metadata is what makes analysis possible, because it turns a store of files and tables into something that can be categorized, compared and trusted.
Why Metadata Creates Risk as Well as Value
Data on its own is inert. Nine unrelated digits can do little damage. Identified as a Social Security number, those same digits become personal information and fall within the scope of state and national data protection law.
Context is what creates the obligation. That makes classification a compliance activity rather than an administrative one, and it means the metadata layer itself has to be governed and protected.
The Scale This Operates At
The World Economic Forum has estimated that more than 463 exabytes of data will be created globally each day by 2025, and put the volume already on the internet above 44 zettabytes. An individual organization generates enough on its own to produce a second category of data describing the first.
Data Governance Is Not a Technology Project
Governance covers availability, usability, integrity and security, and done well it gives teams one place to define and document data and its requirements. It cannot be delivered by tooling alone, because it is a shared responsibility across the organization.
The obstacle is usually perception. Where governance is seen as something that slows the business down, it loses sponsorship and funding, and eventually attracts resentment. Controls that visibly enable work are what prevent that, and each data set should be protected according to its own needs rather than to a single standard applied everywhere.
Governing the Metadata Lifecycle
Data has a lifecycle covering how it is created, managed, monitored, protected and destroyed, and metadata plays a part in every stage of it. The metadata itself then needs a lifecycle of its own, because a catalog that is never reviewed becomes a description of an estate that no longer exists.
Regulation is moving in the same direction. California, Virginia and Colorado have passed state laws, and Canada, Brazil, South Africa, China and the European Union protect their citizens’ data wherever in the world it is held.
Ownership is the practical test of a governance program. Where a data set has a named owner who decides how it may be used and approves changes to that, classification stays accurate as the estate moves around it. Where it does not, the catalog degrades quietly and nobody is accountable for the drift.

