1. Trial
The goal of this trial is to validate the platform’s ability to identify and classify a few key types of sensitive data, demonstrate the resulting reports, and confirm that the insights are meaningful to the client’s compliance and risk-management needs.
Define what “compliance” means to the client
Clarify which framework(s) are most relevant — e.g., GDPR, PCI DSS, HIPAA, or internal data-handling policies.
Emphasize that compliance extends beyond identifying data types; it involves how that data is governed, accessed, and protected across the environment.
Limit the test to 2–3 high-value data patterns
Select a small number of sensitive data types that would be most damaging if exposed (e.g., Social Security numbers, passport numbers, or credit card information).
Keep the focus narrow to ensure the findings are easy to interpret and validate.
Target specific, manageable data locations
Choose one or two data sources that are likely to contain the selected data types. For example:
HR file shares for SSNs or passport numbers (≤ 100 GB total).
Finance folders for credit-card data.
Ask the client where their most sensitive information resides; they usually know or have strong assumptions.Demonstrate discovery and reporting capabilities
Success means the DDC scan runs smoothly, identifies the selected data types in expected locations, and generates clear, actionable reports.
The client should leave the trial with an understanding of:
How the platform classifies and presents sensitive data findings.
What insights and reports they can use for compliance and governance.
Set realistic expectations
Reinforce that this is a proof of concept, not a full implementation.
The aim is not to perfect every classification rule or tune false positives, but to show the visibility, reporting, and control the DDC provides at a small scale.
1.2 Scope the Full Project
Once you’ve completed the trial and decided to proceed, more formal scoping is needed.
Environment mapping
Identify all data stores you want to include: file servers, NAS, cloud storage (OneDrive, SharePoint Online, Dropbox, Exchange Online etc.).
Define components: on-premises servers, cloud tenants, hybrid shares.
Estimate data volume, number of files, number of users, data growth rate.
Define classification model
Define what “sensitive data” means for your organisation (PII, PCI DSS, GDPR etc.).
Create or refine classification patterns, templates, profiles.
Patterns - individual scanning criteria (SSN, UK passport number etc.) - this will be crucial, we must think of what actually is sensitive vs “what could be PII” for example, nobody cares about UK license plate numbers (individually)
Templates - a collection of patterns (custom, or pre-defined) - suggesting to use only custom ones to be sure to know what patterns to scan for
Profiles - this is the actual scanning profile per dataset
Infrastructure
Define required resources: servers, CPU, RAM, storage (for logs, database)
Define database backend (SQL server) and network connectivity (agents to SQL, firewall ports)
- AS PER CURRENT RESOURCE RECOMMENDATIONS PROCESS
Use-case prioritisation and rollout phases
Prioritise which data stores to include first (e.g., high-risk file share used by finance).
Use the information gathered earlier, but dive deeper:
Ask the following questions:
What data sources are in scope? (File servers, SharePoint Online, Exchange Online, OneDrive, NAS, cloud storage, etc.)
How many servers / repositories are there, and where are they located?
What is the estimated total data volume (TB)?
How is data distributed — evenly or concentrated in a few large shares?
Priority list:
Patterns, locations (perhaps there is a company directive, or they’re particularly concerned about sharing sensitive data, or perhaps they’re after a particular compliance (like PCI DSS) for an upcoming audit)
Think about the following:
Server capacity and scalability — you need enough compute for concurrent scanning.
Network throughput — scanning across slow links can bottleneck jobs.
Logical grouping of data sources into manageable “scan sets.”
Tips and Best Practices
Start small: pick a representative segment as a pilot rather than trying to roll out everywhere at once.
Engage business stakeholders early: they understand the data context and can help validate classification results.
Don’t over-classify initially: use high-priority sensitive data types first, then expand.
Monitor system performance: scanning large file volumes can impact servers; schedule scan windows accordingly.
Review false positives: classification rules will need tuning for your environment.
Use classification results to drive action: detection without remediation yields little benefit.
Document processes: classification policy, remediation workflows.
Train users/administrators: make sure staff know what classification results mean and what actions to take.
Integrate with overall security/compliance program: classification alone doesn’t achieve protection unless tied into governance, access control, and monitoring.