EmberHound Personal Data Discovery

Find personal data across every enrolled company device.

EmberHound searches files, documents, and images for personal information, then shows your team where it is stored, what presents the greatest exposure, and where to look when a subject access request arrives. Raw files stay on the device.

  • Windows, macOS, and Linux
  • On-device scanning
  • Raw files are never uploaded
  • Masked findings only

The problem

Personal data spreads through everyday work.

An HR spreadsheet pulled for one review meeting. A customer list exported to help with a mailshot. A scanned sick note saved to a desktop. A support ticket copied into a document so somebody could work on it offline. None of it is anybody being careless - it is just how work gets done.

Interviews, policies, and a manually maintained inventory tell you what your organisation intends to hold. They do not tell you what is on the laptops.

Exports and downloads

Reports pulled for one task, saved to Downloads, and never removed.

Documents and attachments

HR records, scanned forms, screenshots, and email attachments saved locally to work on.

Archives and old folders

Zipped backups and operational folders from projects that ended years ago.

Two jobs

Related, but not the same task.

They use the same findings. They answer different questions, and conflating them is how "GDPR compliance" becomes a word that means nothing.

Personal-data discovery

Ongoing. Understand which categories of personal information exist, where they are stored, and which devices or locations need attention. This is the work that builds the picture.

DSAR search

Reactive, and on a clock. Locate the findings associated with a particular person's identifiers so your team knows where to look. This is the work that answers one request.

What you get

Four things, from the same scan.

See where personal data is stored

Identify supported categories of personal information across enrolled endpoints, by category, device, and location.

Find forgotten or unnecessary copies

Surface data in downloads, exports, archives, and images that will not appear in a manually maintained inventory.

Locate information for a person

Cross-reference known identifiers against existing findings when a subject access request arrives.

Keep evidence of the work

Scan history is retained, and findings and responses export as records of what you searched and what you found.

Detection

What EmberHound looks for.

The core pack covers contact and identity details and government identifiers, and separately the categories that carry extra conditions for processing under the regulation.

Contact and identity

  • Email addresses

    standard address format

  • Telephone numbers

    international dialling formats

  • Postal addresses

    street-suffix patterns, including Strasse, Rue, and Chemin

  • Dates of birth

    day/month/year and ISO forms

  • Names

    matched where a title such as Mr, Ms, Dr, or Prof precedes them

National and government identifiers

  • UK National Insurance numbers

    format-validated

  • UK passport numbers

    nine-digit format

  • UK driving licence numbers

    DVLA format

  • UK Unique Taxpayer Reference

    ten-digit format

  • IBAN

    validated with an IBAN checksum

Article 9

Special category data

Detected by dedicated rules, reported under their own subtype, and switchable per category for your organisation.

Please contact us for details

Article 10

Criminal offence data

Handled separately from Article 9, because it is a different regime with its own conditions for processing rather than a further special category.

Please contact us for details

We do not publish which terms or fields those two sets of rules match on. Setting that out on a public page would tell anybody who wanted to avoid detection how to, and it is not the right place to discuss scanning for information about somebody's health, faith, or criminal record. We will walk through it with you directly.

Two honest limits. Names are matched where a title precedes them, so a bare first and last name in a paragraph will not be picked up on its own. IP addresses and other online identifiers are not detected at all.

Regional detection

National identifiers do not share a format.

A rule that finds a UK National Insurance number will not find a German tax ID. Country packs add the formats and the checksums for a particular country, and every finding records which country rule produced it.

Two country packs are published today. Others are not covered yet, though you can extend detection with custom rule overlays. A country pack tells you what a string is - it does not determine which law applies to your organisation.

United Kingdom

In the core pack

National Insurance number, passport, driving licence, and Unique Taxpayer Reference.

Germany

Country pack

Personalausweis, social security number, and Steuer-ID with its checksum.

France

Country pack

Carte nationale d'identite, NIR with its checksum, and SIREN.

How it works

From agent install to exported evidence.

  1. 01

    Deploy the endpoint agent

    Push it through your existing MDM, or hand somebody an enrolment code.

  2. 02

    Choose what to search for

    Select the personal-data policy, add country packs where relevant, and set which locations are in scope.

  3. 03

    Scan locally

    Matching runs on the device. The file itself never leaves it.

  4. 04

    Review findings

    Category, device, file, masked context, confidence, and risk - grouped and deduplicated.

  5. 05

    Prioritise remediation

    Assign an owner, set a status, work down by risk. Findings carry an SLA due date.

  6. 06

    Rescan and export

    Check whether exposure actually changed, then export the record.

EmberHound never deletes, moves, edits, or redacts the files it scans. Remediation is your team's action, tracked in the platform.

Findings

See enough to investigate without collecting the file.

A finding has to tell somebody where to go and what they will find. It must not become a second copy of the personal data you are trying to get under control. So the preview is masked, the value is stored as a hash, and the rest is metadata.

Findings locate the match inside the file itself: the page of a PDF, the sheet and cell of a spreadsheet, the path within a JSON document.

  • Data category and subtype
  • Masked preview, never the value itself
  • Device, file path, and file type
  • Where inside the file: page, sheet, row, or column
  • Confidence score, and what adjusted it
  • Risk level and exposure level
  • Which country rule applied, where one did
  • How many times it has been seen, across how many scans
  • Status, assigned owner, and resolution note
  • Whether OCR was used to read it

Signal

Focus on credible findings, not every possible match.

A nine-digit number is a passport, a SIREN, an invoice reference, or nothing at all. Six controls separate them, and each finding records which applied to it.

Validation rules

IBAN, the German tax ID, and the French NIR are checksum-validated. Each finding records whether the checksum passed.

Required context

Weak patterns are not reported alone. Special-category rules mostly look for a labelled field and its recorded value, not a word in passing.

Negative matching

The French national ID rule actively excludes invoice, order, ticket, and asset references that share its digit pattern.

Confidence scoring

Each rule carries its own threshold, and every finding records the raw score plus the adjustments applied to it.

Deduplication

A fingerprint groups the same value across scans and devices, so one record in fifty copies of an export reads as one thing.

Suppression rules

Dismiss a finding or exclude a path from future scans. A suppression needs a written reason of at least 10 characters, and it is audited.

None of this eliminates false positives. It moves the credible findings to the top of the list.

Add-on

Find personal data inside images and scanned documents.

A scanned passport, a photographed form, a screenshot of an HR record. Ordinary file search reads none of them, because there is no text in the file to read.

With the OCR add-on enabled, text is extracted from supported images and image-based documents before the same rules run, and each finding records whether OCR produced it.

OCR, mailbox scanning, and drive scanning are add-ons rather than part of a base plan. Ask us about add-on coverage.

  • Applies the same rules to extracted text as to ordinary files
  • Every finding records whether OCR produced it
  • OCR confidence is recorded so low-confidence reads can be filtered
  • Extraction happens on the device, like the rest of scanning

Mailbox scanning is not a cloud integration.

The mailbox add-on extends what the endpoint agent reads. EmberHound does not connect to a mail platform and pull messages from it. If that is what you need, talk to us about your estate before buying.

Subject access requests

When a subject access request arrives, know where to look.

The hard part of a DSAR is rarely the paperwork. It is establishing where a particular person's information actually sits across an estate nobody has a complete picture of. If your endpoints have been scanned, that search is a lookup rather than an expedition.

Identifiers you can search on

  • Email address
  • Telephone number
  • Name
  • Postal address
  • Date of birth
  • National identifier
  • IBAN
  1. 1

    Log the request and verify identity

    Record the request and how identity was confirmed - email confirmation, document upload, knowledge-based checks, a trusted introducer, in person, or account ownership.

  2. 2

    Track the deadline

    The due date is held on the request, along with any extension and the date the subject was told about it.

  3. 3

    Search by identifier

    Identifiers are normalised and hashed with an organisation-specific pepper, then matched against finding hashes. Raw identifiers are not stored in the clear.

  4. 4

    Review what matched

    Matches are clustered and scored into confidence bands. A reviewer includes or excludes each cluster with a reason code, and the decision is recorded against them.

  5. 5

    Produce the response

    An Article 15 response document sets out the personal data, the purposes and legal basis for each record group, an Article 15(1)(c) recipients table, and the subject's rights.

  6. 6

    Deliver and close

    Release through the subject portal, encrypted email, post, or API. The pack records who finalised it, who released it, and when it was delivered.

Where EmberHound stops

It covers the request from intake through to delivery, with audit history throughout. Two things it does not do, deliberately: it does not fetch the underlying documents, because they stay on the endpoint, and it does not redact them. It records the decisions your reviewers make about relevance, exemptions, and third-party information. It does not make those decisions for you.

A list of file locations is not a completed DSAR response. It is the part that used to take the longest.

On the deadline

Organisations normally need to respond without undue delay and within one month. That period can be extended where a request is complex or where somebody has made a number of requests, and the start point can be affected by needing to confirm identity or seek clarification.

It is not a flat 30 days, and it is not the 72 hours that applies to notifying a personal data breach. Those are different obligations with different clocks.

On the search

EmberHound helps your team search supported endpoints consistently and retain evidence of where it looked and what it found. Your organisation remains responsible for deciding the appropriate scope of each request.

A scan is evidence that a search was carried out and how. It is not, by itself, proof that the search was reasonable and proportionate for that particular request.

Records of processing

A record of processing that earns its keep.

The ROPA Workspace holds every Article 30 field, keeps entries under review and sign-off, and feeds signed-off activities into the Article 15 response document. A scan can't tell you your legal basis, your purposes, or your retention period - those are decisions, not facts on a disk. What it can do is show you personal data in places no entry accounts for. Write the record; use the findings to check it against reality.

GDPR

Support GDPR work with evidence from actual data discovery.

Scan results are evidence about storage. That evidence feeds personal-data inventories, data-mapping exercises, retention reviews, exposure remediation, security reviews, internal audits, and the accountability record behind all of them.

How this maps to the Articles

A scanner does not fulfil an Article. It produces evidence that helps you meet an obligation you hold.

Article 5 - Principles

Supports minimisation and accuracy work by showing what is actually stored, and supports accountability by keeping the record of what you found and did.

Article 15 - Right of access

Helps locate information relevant to a request, and produces the response document and delivery record.

Article 16 - Rectification

Supports rectification searches and records the work. EmberHound does not correct the underlying file.

Article 17 - Erasure

Supports erasure-location searches, tracks the work, and can issue a completion certificate. Deleting the file is your action.

Article 30 - Records of processing

Provides the register itself, with the Article 30 fields and a review cycle, plus scan evidence to check the entries against. Completing it does not by itself discharge the obligation.

Article 32 - Security of processing

On-device scanning, masked findings, tenant isolation, and audit history are evidence of technical measures.

EmberHound supports personal-data discovery, DSAR searches, and GDPR accountability work. It does not provide legal advice, guarantee that every relevant record will be found, or make an organisation compliant by itself.

ICO guidance

  1. Right of access: what should we consider when responding to a request? - ICO
  2. Right of access: how do we find and retrieve the relevant information? - ICO

Security

Search personal data without collecting the underlying files.

A tool that found all your personal data by copying it somewhere central would have made the problem worse. Scanning runs on the endpoint and the files stay there.

One exception worth naming: a disclosure pack you choose to produce and deliver does contain the response document your team prepared. That is the point of it, and it is created deliberately rather than as a side effect of scanning.

Read the Trust Centre for the architecture and data-handling detail.

  • Matching runs on the endpoint. The platform never reads your file system.
  • Raw files stay on the device. Only a masked preview, a hash, and metadata leave it.
  • DSAR identifiers are hashed with a pepper held for your organisation alone.
  • Encrypted in transit and at rest.
  • Organisation-level isolation, enforced in the database as well as the application.
  • Role-based access, and every action written to an audit log.

Use cases

What teams run it for.

Map personal data across endpoints

Understand which categories of personal information are actually present on enrolled devices.

Investigate a subject access request

Move from the identifiers you were given to the devices and files your team needs to review.

Find unnecessary local copies

Surface exports, downloads, and archives that no longer have a business purpose behind them.

Review sensitive-category exposure

Article 9 and Article 10 data are covered and gated separately, so each can be looked at on its own terms.

Validate data maps and inventories

Compare what your records say you process against what the devices show you are storing.

Watch whether exposure is improving

Use scan history to see whether data your team dealt with has come back.

Frequently asked questions

Contact and identity details (email, phone, postal address, date of birth, and names carrying a title), UK government identifiers (National Insurance, passport, driving licence, Unique Taxpayer Reference), and IBAN. It also covers Article 9 special category data and, separately, Article 10 criminal offence data - please contact us for details of those. IP addresses and other online identifiers are not detected.

They are covered by dedicated rules, reported under their own subtypes, and switchable per category for your organisation, so you can scan for one without scanning for another. Article 10 criminal offence data is kept separate from the Article 9 categories because it is a different regime with its own conditions for processing, not a further special category. We do not publish the detection detail for either set - please contact us and we will go through it with you. EmberHound identifies these categories; deciding whether you have a lawful basis to hold them is your organisation's judgement.

Enrolled company devices running Windows, macOS, or Linux, and file shares reachable from them. Every finding produced in the platform to date has come from the endpoint agent. Cloud and SaaS connector scanning is not currently available.

No. Matching happens on the device. What leaves it is a masked preview, a fingerprint hash, and the metadata needed to act on the finding - the file path, the category, the confidence score, and where in the file the match sits.

Yes, with the OCR add-on enabled. Text is extracted from supported images and image-based documents before the same rules run, and each finding records whether OCR produced it. Without the add-on, image contents are not read.

Only as an add-on, and not in the way the phrase usually implies. The mailbox add-on extends what the endpoint agent reads. It is not a cloud mailbox integration - EmberHound does not connect to a mail platform and pull messages. If you need that, talk to us about what your estate looks like before buying.

Six controls, and each finding records which of them applied to it. Checksum validation where the identifier supports it, required nearby context for weak patterns, negative matching that actively excludes look-alike references, per-rule confidence thresholds with the adjustments recorded, fingerprint deduplication, and suppression rules for anything reviewed and dismissed. It does not eliminate false positives, and no scanner honestly claims to.

You record the subject's known identifiers. Each is normalised and hashed with a pepper held for your organisation alone, then compared against the hashes stored on findings. Matching happens entirely on hashes - the raw identifier is never stored in the clear. Matches are clustered, scored into confidence bands, and put in front of a reviewer.

It covers the request from intake through to delivery: logging, identity verification, deadline, and extension tracking, identifier search, clustered matching, reviewer inclusion and exclusion decisions, the Article 15 response document, portability exports, erasure certificates, delivery, and closure - all with audit history. What it does not do is fetch the underlying documents or decide for you what is relevant, exempt, or third-party information. Those judgements stay with your team.

No. Findings show masked previews rather than values, and a reviewer can exclude a cluster from a disclosure with a reason. But EmberHound does not open your source documents and redact them - it never modifies the files it scans.

No. Rectification and erasure are supported as workflows: the request is recorded, the relevant locations are identified, the work is tracked, and an erasure certificate can be issued on completion. Deleting or correcting the underlying file is an action your team takes.

It gives you the register and the fields - purpose, controller, legal basis, categories of personal data, categories of data subjects, recipients, retention period, transfer countries and safeguards, and security measures - along with a draft, needs-review, complete lifecycle and a next review date. What it does not do is write the entries for you or derive them from a scan: a scan cannot tell you your legal basis or your retention period. It also does not judge whether an entry is adequate. Your team authors the record; the findings help you check it against what is actually stored.

No. It supports personal-data discovery, DSAR searches, and accountability work by producing evidence about where data is stored and what was done about it. Compliance is a judgement your organisation makes, with advice where it needs it.

Organisations normally need to respond without undue delay and within one month. The period can be extended where a request is complex or where you have received a number of requests from the same person, and the start point can be affected by needing to confirm identity or seek clarification. The ICO's guidance is the authority on how those situations work - we link to it below rather than reproducing it.

One device and one user, with a 500 GB monthly scanning limit and one scan, and no card details required. It covers contact details, UK identifiers, and IBAN. The findings dashboard, masked previews, and scan history work as they do on paid plans. Special category data, OCR, and DSAR search need a paid plan.

OCR for image-based documents, mailbox scanning, and cloud or external drive scanning are add-ons on top of any plan. Personal-data discovery and the DSAR workflow are part of the plans that include them - see the pricing page for what each one covers.

Find out where personal data is being stored.

Run your first scan on one enrolled device and see the categories, locations, and exposure EmberHound identifies, without uploading the underlying files.

No card details required. One device, one user.

Your cookie choices

We use cookies to run this site, measure how it is used, and to advertise on other platforms. You can accept or refuse each purpose separately.

Keeps you signed in and remembers this choice. Always on.

Google Analytics, Sentry and Vercel. Which pages are used, and what breaks.

LinkedIn, X and Meta pixels, loaded through Google Tag Manager.

Cookie policy