Security fundamentals

What is data classification?

Data classification is the practice of sorting data into sensitivity levels, commonly Public, Internal, Confidential, and Restricted, so you can apply the right protection to each. It is the foundation for access control, encryption, and retention decisions.

Definition

Data classification is the process of categorizing information by how sensitive it is, so that the right handling, access, and protection rules can be applied consistently to each category.

Background

Most schemes use a small number of tiers, for example Public, Internal, Confidential, and Restricted. Each tier carries handling rules: who may access it, whether it must be encrypted, how it may be shared, and how long it is retained. Classification underpins other controls, because you cannot apply least privilege, encryption, or retention sensibly until you know how sensitive the data is. SOC 2 and ISO 27001 both expect a classification scheme.

Why it matters

Without classification, everything is treated the same, which means either over-protecting low-risk data (friction) or under-protecting sensitive data (risk). A simple, well-understood scheme lets you focus effort where it matters and prove you handle sensitive data appropriately.

Step by step

  1. Define a small number of clear levels (three or four is usually enough).
  2. Write handling rules for each level: access, encryption, sharing, and retention.
  3. Inventory your key data stores and assign each a classification.
  4. Label data and systems so the classification is visible and actionable.
  5. Train people on what the levels mean and how to handle each.
  6. Review classifications as data and systems change.

Examples

  • Customer records and secrets are classified Restricted, encrypted, and limited to a few roles; the public marketing site is Public.
  • An internal wiki is Internal: accessible to employees but not shared externally without review.

Common mistakes

  • Creating too many levels, so no one can remember or apply them.
  • Classifying data but never writing the handling rules that make it useful.
  • Never revisiting classifications as new systems and data appear.

FAQ

What are common data classification levels?

A typical scheme uses Public, Internal, Confidential, and Restricted, though the exact labels vary. The key is a small, clear set of tiers each with defined handling rules.

Why does data classification matter for compliance?

It is the basis for applying access control, encryption, and retention proportionately, and SOC 2 and ISO 27001 expect you to have a scheme and follow it.

Related

What is least privilege? → What is encryption? → Policy management in Keel →

Do this in Keel, not a spreadsheet

Keel is the AI-native GRC platform for SMBs: one control-and-evidence graph across SOC 2, ISO 27001, HIPAA, PCI DSS, NIST CSF, and more. Start free, no credit card.

Start free