Shadow AI is already using your data. Get the complimentary Gartner® report. Read the report

Data Privacy and Protection at Global Biotech

PKWARE

By PKWAREProductivity Protected

Share on social media

Company Profile

Company
Global Biotech
Size
Large Enterprise
Industry
Healthcare
Location
California, USA

A California biotechnology company with about 15,000 employees was building a data lake for analytics and needed everything entering it de-identified first. Seventeen classifiers, third-party JSON and CSV feeds, and a fixed ingestion window left no room for manual review. PK Protect combined custom classification rules, API-driven automation and EMR-scale processing, so discovery and masking ran inside the ingestion flow rather than after it.

Background

This PKWARE customer is a biotechnology company based in California and is dedicated to scientific research and development for medicines to treat people with serious and life-threatening diseases. The company employs almost 15,000 employees and achieves nearly $2 billion in annual revenue.

Two figures set the constraints. Fifteen thousand employees generate data across research, clinical and commercial functions, and roughly two billion dollars in revenue funds an analytics program that expects that data to be usable. Protection therefore could not mean withholding data from the teams asking for it.

Challenges

  1. Different customer projects were using a variety of data types. They needed a solution that offered flexible policy creation to address finding different kinds of sensitive data ranging from PHI to HIPAA data. They had 17 different classifiers. Every data set would need a new policy.
  2. The customer was setting up a data lake into which they were going to ingest data from a variety of third-party data sets coming primarily in JSON and CSV files.
  3. PKWARE needed to work with AWS as well as with a designated solution provider. The PS Consultant was stationed on site with the customer. The proof of concept (POC) was partially completed with AWS. The next step was to have the additional solution provider independently test PKWARE APIs and automation. The final steps would include fine-tuning the rules and creating a de-identified data lake with fully protected data to be consumed by the customer’s analytics team.

Key requirements included:

  • Performance was key. They also had strict time restrictions for data ingestion.
  • A hands-off, automated solution was extremely important.
  • Accuracy in minimizing false positives in an unstructured format was also crucial.

The seventeen classifiers are the difficult part rather than the volume. A policy written per data set multiplies with every new project, and each copy is a place where a rule can drift out of alignment with the rest. Flexible policy creation matters because one rule set can then cover many data sets, which is what keeps classification consistent as the number of projects grows.

Our Approach

Using PKWARE, every element of sensitive data was identified and masked as required for compliance.

Download PDF

Use Cases

Step 1:

  • Leveraged certain out-of-the-box elements from the PKWARE platform. Then used the PKWARE policy builder to build custom rules (elements) for their unique, specific elements: patient IDs and serial numbers.
  • Fine-tuned rules to capture the nuances of their environment, resulting in no false positives and no false negatives.

Step 2:

  • Created an automated flow. Leveraged PKWARE APIs and AWS Lambda functions.

Step 3:

  • Leveraged EMR at scale to achieve detection and masking for a specific window of time.

JSON and CSV arriving from third parties explain why tuning came before automation. Semi-structured files do not guarantee that a given field sits in a given place, so classification has to work from content rather than position. Removing false positives first is what makes the later automated stages safe, because an automated masking flow acts on whatever the classifier returns with no human in the loop.

Results

Fully Automated Solution Processes that had been manual, slow, and inaccurate became rapid and accurate.

Full De-Identification Every element of sensitive data was identified and masked as required for compliance.

Safe Zone for Analytics Fully protected data now flows into the customer’s safe zone data lake. Analytics can deliver full value safely.

De-identification rather than restriction is the outcome worth noting. Masked data still supports reporting and modelling, so the safe zone is a working analytics environment rather than an archive. The processes that were manual and slow were also the ones that were inaccurate, which is the usual pattern: the review that takes longest is the one most likely to miss something.

PKWARE

PKWARE

Productivity Protected

PKWARE has been securing sensitive data for over 40 years. We’ve earned the trust of 21 of the 25 largest banks in the U.S. Our team delivers modern, data-centric security solutions organizations can rely on.