How to Trial and Deploy our DDC

1. Trial

The goal of this trial is to validate the platform’s ability to identify and classify a few key types of sensitive data, demonstrate the resulting reports, and confirm that the insights are meaningful to the client’s compliance and risk-management needs.

  • Define what “compliance” means to the client

  • Clarify which framework(s) are most relevant — e.g., GDPR, PCI DSS, HIPAA, or internal data-handling policies.

  • Emphasize that compliance extends beyond identifying data types; it involves how that data is governed, accessed, and protected across the environment.

  • Limit the test to 2–3 high-value data patterns

  • Select a small number of sensitive data types that would be most damaging if exposed (e.g., Social Security numbers, passport numbers, or credit card information).

  • Keep the focus narrow to ensure the findings are easy to interpret and validate.

  • Target specific, manageable data locations

  • Choose one or two data sources that are likely to contain the selected data types. For example:

  • HR file shares for SSNs or passport numbers (≤ 100 GB total).

  • Finance folders for credit-card data.
    Ask the client where their most sensitive information resides; they usually know or have strong assumptions.

  • Demonstrate discovery and reporting capabilities

  • Success means the DDC scan runs smoothly, identifies the selected data types in expected locations, and generates clear, actionable reports.

  • The client should leave the trial with an understanding of:

  • How the platform classifies and presents sensitive data findings.

  • What insights and reports they can use for compliance and governance.

  • Set realistic expectations

  • Reinforce that this is a proof of concept, not a full implementation.

  • The aim is not to perfect every classification rule or tune false positives, but to show the visibility, reporting, and control the DDC provides at a small scale.

1.2 Scope the Full Project

Once you’ve completed the trial and decided to proceed, more formal scoping is needed.

  • Environment mapping

    • Identify all data stores you want to include: file servers, NAS, cloud storage (OneDrive, SharePoint Online, Dropbox, Exchange Online etc.).

    • Define components: on-premises servers, cloud tenants, hybrid shares.

    • Estimate data volume, number of files, number of users, data growth rate.

  • Define classification model

    • Define what “sensitive data” means for your organisation (PII, PCI DSS, GDPR etc.).

    • Create or refine classification patterns, templates, profiles.

Patterns - individual scanning criteria (SSN, UK passport number etc.) - this will be crucial, we must think of what actually is sensitive vs “what could be PII” for example, nobody cares about UK license plate numbers (individually)

Templates - a collection of patterns (custom, or pre-defined) - suggesting to use only custom ones to be sure to know what patterns to scan for

Profiles - this is the actual scanning profile per dataset

  • Infrastructure

    • Define required resources: servers, CPU, RAM, storage (for logs, database)

    • Define database backend (SQL server) and network connectivity (agents to SQL, firewall ports)

 - AS PER CURRENT RESOURCE RECOMMENDATIONS PROCESS

  • Use-case prioritisation and rollout phases

    • Prioritise which data stores to include first (e.g., high-risk file share used by finance).

    • Use the information gathered earlier, but dive deeper:

  • Ask the following questions:

    • What data sources are in scope? (File servers, SharePoint Online, Exchange Online, OneDrive, NAS, cloud storage, etc.)

    • How many servers / repositories are there, and where are they located?

    • What is the estimated total data volume (TB)?

    • How is data distributed — evenly or concentrated in a few large shares?

Priority list:

Patterns, locations (perhaps there is a company directive, or they’re particularly concerned about sharing sensitive data, or perhaps they’re after a particular compliance (like PCI DSS) for an upcoming audit)

  • Think about the following:

    • Server capacity and scalability — you need enough compute for concurrent scanning.

    • Network throughput — scanning across slow links can bottleneck jobs.

    • Logical grouping of data sources into manageable “scan sets.”

Tips and Best Practices

  • Start small: pick a representative segment as a pilot rather than trying to roll out everywhere at once.

  • Engage business stakeholders early: they understand the data context and can help validate classification results.

  • Don’t over-classify initially: use high-priority sensitive data types first, then expand.

  • Monitor system performance: scanning large file volumes can impact servers; schedule scan windows accordingly.

  • Review false positives: classification rules will need tuning for your environment.

  • Use classification results to drive action: detection without remediation yields little benefit.

  • Document processes: classification policy, remediation workflows.

  • Train users/administrators: make sure staff know what classification results mean and what actions to take.

  • Integrate with overall security/compliance program: classification alone doesn’t achieve protection unless tied into governance, access control, and monitoring.