
Reviews
Personal data, sensitive data, metadata, inferred data, anonymized data, and pseudonymous data compared
Personal data categories compared by identifiability, sensitivity, source, inference, reversibility, legal status, control needs, and practical privacy risk.
What to take away
- These labels overlap; one record can be personal, sensitive, inferred, and metadata at once.
- Identifiability depends on context and reasonably available combinations, not names alone.
- Sensitive data describes higher-impact content or use, with definitions that vary by law.
- Metadata can expose relationships, timing, movement, and behavior.
- Pseudonymized data remains linkable through additional information.
- Anonymization requires a defensible assessment of re-identification, not simple redaction.
Data categories help people choose controls, but labels can mislead when treated as sealed boxes. A timestamp may look harmless by itself, yet a series of locations and times can identify a routine. A coded research record may omit a name while remaining linkable through a separate key.
Side-by-side view
Data categories compared
Category
- Personal data
- Identifiable person?
- Sensitive data
- Heightened harm?
- Metadata
- Describes event/object?
- Inferred data
- Conclusion produced?
- Pseudonymous data
- Separate route back?
- Anonymized data
- No reasonable identification?
Basic question
- Personal data
- Purpose and rights
- Sensitive data
- Strict access, necessity
- Metadata
- Pattern exposure
- Inferred data
- Accuracy, decision effects
- Pseudonymous data
- Separation, key control
- Anonymized data
- Re-identification assessment
Main control concern
- Personal data
- Sensitive data
- Metadata
- Inferred data
- Pseudonymous data
- Anonymized data
Basic question
- Personal data
- Can it relate to an identifiable person?
- Sensitive data
- Could content or use create heightened harm?
- Metadata
- What describes an event or object?
- Inferred data
- What conclusion was produced?
- Pseudonymous data
- Is a separate route back retained?
- Anonymized data
- Is identification no longer reasonably possible in context?
Example
- Personal data
- Account email
- Sensitive data
- Precise location
- Metadata
- Message time and recipient
- Inferred data
- Predicted interest
- Pseudonymous data
- Coded participant ID
- Anonymized data
- Properly aggregated statistics
Main control concern
- Personal data
- Purpose and rights
- Sensitive data
- Strict access and necessity
- Metadata
- Relationship and pattern exposure
- Inferred data
- Accuracy and decision effects
- Pseudonymous data
- Separation and key control
- Anonymized data
- Re-identification assessment
Personal data
Personal data is broader than a name. It can include a direct identifier or information that identifies a person indirectly when combined with other available material. Account IDs, device identifiers, photographs, voices, locations, employment details, and online behavior may qualify depending on context and law.
The ICO's UK-focused explanation of what counts as personal data emphasizes identified or identifiable living individuals and warns that information labeled anonymous may remain personal if reasonably available means can reconnect it. The page supports that narrow UK GDPR point, not a worldwide definition.
Sensitive data
"Sensitive" can be a risk description, a contract label, or a defined legal class. Common examples include health, biometrics used for identification, financial credentials, precise location, intimate communications, government identifiers, and information about children. Statutes use different names and boundaries.
Do not rely on the category alone. A home address may be ordinary shipping data in one context and dangerous in a stalking case. Evaluate severity, likelihood, volume, persistence, recipients, and the person's circumstances.
Metadata
Metadata is data about another item or event. An email can carry sender, recipient, subject, time, routing, and attachment details. A photograph can include capture time, device information, and coordinates. A purchase can produce merchant, amount, time, and device records.
Protecting message content while exposing a complete contact graph may still reveal sensitive relationships. Remove unnecessary metadata before publishing files and restrict access to operational logs.
Inferred data
Inferences are conclusions produced from observed, supplied, purchased, or modeled inputs. They include likely interests, fraud risk, household composition, purchasing power, health tendencies, or account authenticity.
An inference can be wrong and still affect a person. Record its source, purpose, confidence, refresh cycle, review path, and consequences. A prediction should not be presented as a confirmed fact.
Pseudonymous data
Pseudonymization replaces or separates identifying elements so the record cannot be attributed without additional information. The additional key or lookup route remains important. If the same organization can reconnect the pieces, it still handles personal data.
The ICO's detailed pseudonymisation guidance explains that the pseudonymized dataset and additional information can reconstruct the original personal data and should be kept separate with suitable controls. It is jurisdiction-specific guidance, not a claim that one coding method is sufficient everywhere.
Anonymized data
Anonymization aims to make people no longer identifiable under the applicable test and context. Removing names is only one step. Rare attributes, exact dates, geography, external datasets, repeated identifiers, and small groups can allow singling out or reconnection.
Assess the data, likely recipients, available outside information, technical effort, incentives, and future change. Controls on access and attempted re-identification can reduce risk, but a contractual promise alone does not transform identifiable records into anonymous information.
Choose controls by the real risk
Use overlapping labels when needed. A voice-message timestamp may be personal metadata. A predicted medical condition can be personal, sensitive, and inferred. A coded genetics record can be sensitive and pseudonymous.
For each set, document:
- identifiability and possible combinations;
- sensitivity and plausible harm;
- source and accuracy;
- decision or service purpose;
- recipients and access;
- linkability and retained keys;
- retention and deletion;
- rights, corrections, and challenges.
The control follows the record's actual use and context, not the least demanding label available.
Common questions
Is an IP address always personal data?
The answer depends on applicable law and whether the address can relate to an identifiable person in context. Avoid a universal claim.
Is aggregated data automatically anonymous?
No. Small groups, rare combinations, repeated releases, and outside datasets can permit singling out or re-identification.
Does hashing make an identifier anonymous?
Not automatically. Repeated hashes can remain linkable, and predictable inputs may be tested. Context and retained information matter.
Can metadata be more revealing than content?
Sometimes. A pattern of recipients, locations, and times can reveal relationships or routines even when content is protected.





