Comparison card of personal, sensitive, metadata, inferred, anonymized, pseudonymous data categories. Personal data, sensitive data, metadata, inferred data, anonymized data, and pseudonymous data compared
Image: Privacy Notes

Reviews

Part of Personal data privacy guide: collection, use, sharing, retention, security, deletion, and accountability

Personal data, sensitive data, metadata, inferred data, anonymized data, and pseudonymous data compared

Personal data categories compared by identifiability, sensitivity, source, inference, reversibility, legal status, control needs, and practical privacy risk.

What to take away

  • These labels overlap; one record can be personal, sensitive, inferred, and metadata at once.
  • Identifiability depends on context and reasonably available combinations, not names alone.
  • Sensitive data describes higher-impact content or use, with definitions that vary by law.
  • Metadata can expose relationships, timing, movement, and behavior.
  • Pseudonymized data remains linkable through additional information.
  • Anonymization requires a defensible assessment of re-identification, not simple redaction.

Data categories help people choose controls, but labels can mislead when treated as sealed boxes. A timestamp may look harmless by itself, yet a series of locations and times can identify a routine. A coded research record may omit a name while remaining linkable through a separate key.

Side-by-side view

Data categories compared

Category

Personal data
Identifiable person?
Sensitive data
Heightened harm?
Metadata
Describes event/object?
Inferred data
Conclusion produced?
Pseudonymous data
Separate route back?
Anonymized data
No reasonable identification?

Basic question

Personal data
Purpose and rights
Sensitive data
Strict access, necessity
Metadata
Pattern exposure
Inferred data
Accuracy, decision effects
Pseudonymous data
Separation, key control
Anonymized data
Re-identification assessment

Main control concern

Personal data
Sensitive data
Metadata
Inferred data
Pseudonymous data
Anonymized data

Basic question

Personal data
Can it relate to an identifiable person?
Sensitive data
Could content or use create heightened harm?
Metadata
What describes an event or object?
Inferred data
What conclusion was produced?
Pseudonymous data
Is a separate route back retained?
Anonymized data
Is identification no longer reasonably possible in context?

Example

Personal data
Account email
Sensitive data
Precise location
Metadata
Message time and recipient
Inferred data
Predicted interest
Pseudonymous data
Coded participant ID
Anonymized data
Properly aggregated statistics

Main control concern

Personal data
Purpose and rights
Sensitive data
Strict access and necessity
Metadata
Relationship and pattern exposure
Inferred data
Accuracy and decision effects
Pseudonymous data
Separation and key control
Anonymized data
Re-identification assessment

Personal data

Personal data is broader than a name. It can include a direct identifier or information that identifies a person indirectly when combined with other available material. Account IDs, device identifiers, photographs, voices, locations, employment details, and online behavior may qualify depending on context and law.

The ICO's UK-focused explanation of what counts as personal data emphasizes identified or identifiable living individuals and warns that information labeled anonymous may remain personal if reasonably available means can reconnect it. The page supports that narrow UK GDPR point, not a worldwide definition.

Sensitive data

"Sensitive" can be a risk description, a contract label, or a defined legal class. Common examples include health, biometrics used for identification, financial credentials, precise location, intimate communications, government identifiers, and information about children. Statutes use different names and boundaries.

Do not rely on the category alone. A home address may be ordinary shipping data in one context and dangerous in a stalking case. Evaluate severity, likelihood, volume, persistence, recipients, and the person's circumstances.

Metadata

Metadata is data about another item or event. An email can carry sender, recipient, subject, time, routing, and attachment details. A photograph can include capture time, device information, and coordinates. A purchase can produce merchant, amount, time, and device records.

Protecting message content while exposing a complete contact graph may still reveal sensitive relationships. Remove unnecessary metadata before publishing files and restrict access to operational logs.

Inferred data

Inferences are conclusions produced from observed, supplied, purchased, or modeled inputs. They include likely interests, fraud risk, household composition, purchasing power, health tendencies, or account authenticity.

An inference can be wrong and still affect a person. Record its source, purpose, confidence, refresh cycle, review path, and consequences. A prediction should not be presented as a confirmed fact.

Pseudonymous data

Pseudonymization replaces or separates identifying elements so the record cannot be attributed without additional information. The additional key or lookup route remains important. If the same organization can reconnect the pieces, it still handles personal data.

The ICO's detailed pseudonymisation guidance explains that the pseudonymized dataset and additional information can reconstruct the original personal data and should be kept separate with suitable controls. It is jurisdiction-specific guidance, not a claim that one coding method is sufficient everywhere.

Anonymized data

Anonymization aims to make people no longer identifiable under the applicable test and context. Removing names is only one step. Rare attributes, exact dates, geography, external datasets, repeated identifiers, and small groups can allow singling out or reconnection.

Assess the data, likely recipients, available outside information, technical effort, incentives, and future change. Controls on access and attempted re-identification can reduce risk, but a contractual promise alone does not transform identifiable records into anonymous information.

Choose controls by the real risk

Use overlapping labels when needed. A voice-message timestamp may be personal metadata. A predicted medical condition can be personal, sensitive, and inferred. A coded genetics record can be sensitive and pseudonymous.

For each set, document:

  • identifiability and possible combinations;
  • sensitivity and plausible harm;
  • source and accuracy;
  • decision or service purpose;
  • recipients and access;
  • linkability and retained keys;
  • retention and deletion;
  • rights, corrections, and challenges.

The control follows the record's actual use and context, not the least demanding label available.

Common questions

Is an IP address always personal data?

The answer depends on applicable law and whether the address can relate to an identifiable person in context. Avoid a universal claim.

Is aggregated data automatically anonymous?

No. Small groups, rare combinations, repeated releases, and outside datasets can permit singling out or re-identification.

Does hashing make an identifier anonymous?

Not automatically. Repeated hashes can remain linkable, and predictable inputs may be tested. Context and retained information matter.

Can metadata be more revealing than content?

Sometimes. A pattern of recipients, locations, and times can reveal relationships or routines even when content is protected.

More in Reviews

Latest from Costs Desk