Last Updated on August 3, 2026 by Satyendra
Sensitive data rarely stays where you expect it to. A Social Security number is stored in a file shared in a spreadsheet. A patient record is sent as an attachment to an email. Someone’s laptop has a report that contains a card number. Now multiply that across thousands of employees and years’ worth of stored files. Most organizations end up with financial data, PHI, and PII dispersed throughout that no one is monitoring.
You can’t protect what you can’t see, and regulators no longer accept “we didn’t know it was there” as an excuse.
This guide explains what these data types are, why it’s important to identify them for compliance, where they usually hide, the challenges involved, and how to detect and classify them manually or with automated methods.
What is PII, PHI, and Financial Data?
Before you can identify sensitive data, you need a clear understanding of what counts as such.
- Personally Identifiable Information (PII) includes an individual’s full name, Social Security number, driver’s license, mailing address, passport information, financial information, email address, and phone number.
- Protected Health Information (PHI) includes medical record numbers and diagnosis codes, insurance and billing information, treatment history, and prescription details.
- Financial Data includes details about a person’s or an organization’s financial accounts and transactions, including bank account information, credit and debit card numbers, payroll and salary information, transaction history, and invoices.
Exposure of PII, PHI, or financial data can lead to identity theft, financial fraud, phishing, and insurance fraud, causing serious and long-term harm. As a result, regulations govern how this data is collected, stored, processed, and protected, including:
- GDPR (General Data Protection Regulation) governs personal data of EU residents and requires organizations to know where that data lives and how it is processed.
- HIPAA (Health Insurance Portability and Accountability Act) requires the U.S. healthcare sector to protect PHI.
- PCI DSS (Payment Card Industry Data Security Standard) establishes guidelines for businesses that handle, store, and transmit cardholder data
Why Locating Sensitive Files Matters for Compliance
Locating sensitive data is a foundational step for meeting modern data privacy and security regulations.
| Reason | Explanation |
|---|---|
| Regulatory Requirements | Regulations such as GDPR, CCPA, and HIPAA require organizations to identify, inventory, and protect sensitive data. |
| Risk of Data Breaches and Insider Threats | Untracked or unclassified sensitive files are easier for attackers and insiders to misuse without detection. |
| Audit and Report Challenges | Without knowing where sensitive data resides and who can access it, organizations face time-consuming audits and potential compliance gaps. |
| Business Impact | Sensitive data breaches can lead to regulatory fines, reputational damage, loss of customer trust, legal liability, and costly remediation. |
How to Identify Files Containing PII, PHI, or Financial Data across File Systems and Cloud
On Windows File Servers
Windows Server includes File Server Resource Manager (FSRM), which provides basic file classification capabilities. Administrators can create classification rules to identify files containing PII, PHI, or financial data based on file properties, file location, or content using classification methods such as keywords or regular expressions. While FSRM is not a comprehensive data discovery platform, it can help locate regulated data on Windows file servers.
Steps to Identify Sensitive Files Using FSRM
- Install the FSRM role through Server Manager, if required.
- Open FSRM and go to Classification Management.
- Create Classification Properties and Rules to scan specific folders or file shares.
- Configure rules to detect identifiers such as Social Security numbers, passport numbers, credit card numbers, or bank account numbers.
- Run classification manually or schedule it to run automatically.
- Review and export the classification results to identify files containing PII, PHI, or financial data. Then, review file permissions separately to determine whether additional security controls are required.
FSRM is useful for identifying regulated data on Windows file servers but relies on manually configured rules and provides limited visibility beyond on-premises file systems.
On Microsoft 365
Microsoft 365 provides data discovery capabilities through Microsoft Purview. Using Sensitive Information Types, Content Explorer, and Data Loss Prevention (DLP), organizations can identify files containing PII, PHI, or financial data across services such as SharePoint Online, OneDrive, Exchange Online, and Microsoft Teams.
Steps to Identify Sensitive Files
- Sign in to the Microsoft Purview portal with the required permissions.
- Review built-in Sensitive Information Types for data such as Social Security numbers, passport numbers, credit card numbers, bank account numbers, and medical record numbers, health insurance claim numbers, or National Provider Identifiers (NPIs). Create custom types if needed.
- Configure DLP or other Purview policies using the required sensitive information types
- Allow Microsoft Purview to index and evaluate supported Microsoft 365 workloads based on the configured sensitive information types and policies.
- Use Content Explorer (or Data Explorer, if available) to review locations containing files in which Microsoft Purview has detected sensitive information
- Filter results by data type, workload, location, or sensitivity label.
- Apply appropriate remediation, such as sensitivity labels, retention labels, DLP policies, or access permission changes where appropriate.
Microsoft Purview provides broader discovery across Microsoft 365 than manual searches, but advanced discovery and reporting features depend on the applicable Microsoft 365 licensing.
Enterprise Workflow for Sensitive Data Discovery
Regardless of the tools used, organizations should follow a structured workflow to ensure sensitive data discovery remains accurate and repeatable.
- Define the Discovery Scope: Identify which types of sensitive information need to be located, such as PII, PHI, financial records, or intellectual property. Determine which repositories, departments, and business units are within scope.
- Inventory Data Repositories: Document all locations where sensitive information may exist, including Windows file servers, NAS devices, SharePoint Online, OneDrive, Exchange Online, cloud storage services, endpoints, and SaaS applications
- Configure Detection Policies: Define the sensitive information types, keywords, regular expressions, confidence levels, and classification rules that will be used during scanning.
- Run Discovery Scans: Scan all repositories within scope to identify files containing sensitive information. Initial scans establish a baseline, while scheduled scans identify newly created or modified files.
- Review and Validate Findings: Confirm that detected files actually contain sensitive information and adjust detection rules where necessary to reduce false positives and improve accuracy.
- Prioritize Remediation: Determine which files present the highest risk based on their sensitivity, location, sharing status, and user permissions. Apply appropriate security controls such as access reviews, sensitivity labels, encryption, or retention policies
- Continuously Monitor Sensitive Data: Sensitive data discovery should be an ongoing process rather than a one-time activity. Continuous monitoring helps detect newly created sensitive files, permission changes, and changes in exposure over time
Finding Sensitive Data is Only Half the Problem
Knowing where PII, PHI, and financial data reside is important, but location alone does not determine risk. Organizations must also know who can access it and whether they still need that access. A sensitive file that is accessible to one authorized user poses less risk than one exposed to hundreds through inherited permissions, open shares, or stale accounts. Therefore, effective sensitive data protection requires linking data discovery with identity and access visibility.
Challenges and Limitations of Native Data Discovery Tools
Finding sensitive data across modern IT environments is challenging, even with native discovery tools. Organizations often face:
- Limited Visibility Across Hybrid Environments: Data spread across on-premises file servers, cloud platforms, and SaaS applications often requires multiple discovery solutions to maintain complete visibility.
- Unstructured and Inaccessible Data: PDFs, spreadsheets, Word Documents, scanned images, encrypted files, and compressed archives can make content inspection difficult.
- Manual Classification Rules: Native tools may require administrators to create and maintain keywords, classification rules, and regular expressions as data and regulatory requirements evolve.
- Data Sprawl and Redundancy: Duplicate files across multiple locations increase the volume of data that must be identified, reviewed, and protected.
- Limited Permission Correlation: Identifying sensitive files is only part of the challenge. Native tools may provide limited insight into who can access them, making actual exposure difficult to assess.
- Licensing Considerations: Some advanced Microsoft Purview discovery and classification capabilities require higher-tier Microsoft 365 licenses.
- Fragmented Reporting: Reporting across separate administrative consoles can make it difficult to gain a consolidated view of sensitive data, permission changes, file activity, and ongoing exposure risks.
How to Choose the Right Data Classification Tool
When evaluating sensitive data discovery and classification solutions, make sure they cover these aspects:
- Accurate Detection: Quickly identifies sensitive data such as PII, PHI, and financial information while minimizing false positives.
- Cloud and On-Premises Coverage: A single solution should be able to find sensitive data across different infrastructures such as file servers, NAS devices, cloud storage, and SaaS platforms.
- Real-Time Monitoring: Instead of waiting for scheduled scans, real-time monitoring can detect when data is created, changed, or moved.
- Reporting and Compliance: Provides ready-to-use reports for regulations such as GDPR, HIPAA, and PCI DSS, making audits easier.
- Integration with Access and DLP Tools: Data discovery results can be integrated with your access control and DLP tools to help ensure that sensitive data is adequately protected
- Easy to Deploy and Scale: Deploy and scale as your organization’s data and storage needs grow.
How Lepide Helps Find and Classify Sensitive Data
Lepide’s data discovery and classification solution empowers organizations by revealing the location of their sensitive data within the environment and continuing to keep track of the changing data locations as environments evolve.
- Automated Discovery across File Systems and Cloud: Scan file servers, NAS devices, and cloud platforms such as OneDrive, SharePoint, and Google Drive from a single console to locate sensitive data
- Built-in Classification Rules: Use pre-configured regulatory templates to simplify setup and improve sensitive data detection.
- Real-Time Alerts on Sensitive Data Exposure: Get notified about sensitive file activities, such as creation, movement, or access, to respond quickly to potential risks.
- Visibility into Who Has Access to What: Identify where sensitive data is stored and who can access it, helping detect unnecessary or risky permissions.
- Simplify Compliance Reporting: Generate reports to support compliance with regulations such as GDPR, HIPAA, and PCI DSS and streamline audit preparation.
Schedule a demo today and see how Lepide automatically discovers PII, PHI, and financial data across your systems, before it becomes a breach!
Conclusion
You can’t protect sensitive data if you don’t know where it exists. PII, PHI, and financial data can quickly spread across file servers, cloud platforms, emails, endpoints, and SaaS applications, making manual tracking difficult. Regular, automated data discovery and classification help organizations locate sensitive data and identify potential exposure before it leads to a breach or compliance violation.
Data discovery is not just about compliance; it is the foundation of effective data security. Access controls, DLP, encryption, and monitoring work more effectively when organizations know where sensitive data is stored, who can access it, and how it is being used.
Frequently Asked Questions
Detection of sensitive content through modern tools can happen among various types of files, including document files, spreadsheet files, PDFs, and emails. It is also possible to scan file servers, NAS devices, cloud platforms like OneDrive, SharePoint, and Google Drive, apart from some SaaS applications.
Sensitive data discovery must be conducted on a continuous basis or on set intervals at a minimum. Given that new files are being created and modified on a day-to-day basis, it helps if scans are done regularly so that newly exposed sensitive data can readily be spotted by the companies. Annual or quarterly audits alone may not be enough.
Organizations use data discovery and classification tools to find sensitive information such as PII, PHI, and financial data. These tools use pattern matching, regular expressions, and AI/ML used by such tools to automatically flag identified files or data records. Tools collaborate with access governance and DLP software to ensure this information remains protected.
By optimizing detection rules and regular expression patterns, organizations can reduce false positives or incorrect detection of confidential data during the scan.
In place of pure pattern matching, AI/ML classification methods that are context-sensitive will minimize errors, and finally, reviewing periodic scan results helps refine detection rules and reduce incorrect flagging.