← All articles

PRIVIOT BLOG

What data can a PDF contain? Metadata, attachments, and hidden text explained

PDFs can contain more than visible pages. Learn how metadata, comments, attachments, form data, links, and OCR text can remain inside a file.

A PDF often feels like a finished digital sheet of paper. Technically, it can contain much more than what you see on the page. Text and images may sit alongside metadata, comments, form values, attachments, links, and invisible text layers.

None of that is inherently bad. Many of these features are useful. They matter when you publish a document, anonymize it, or share it with people who should not receive every piece of information stored in the file.

Visible page content

The obvious part of a PDF is its pages: text, images, lines, shapes, and other graphical objects. Even here, the internal structure can vary. A paragraph may be stored as real text, while a scanned letter may be nothing more than an image.

That is why PDFs that look similar can behave very differently when you try to search or edit them.

Metadata

Many PDFs contain document properties such as title, author, subject, keywords, and creation or modification information. The exact fields depend on the software and workflow used to create the file.

Metadata can be useful for archiving and document management. But if you are publishing or anonymizing a document, it may reveal information that does not appear on the visible pages.

Comments and annotations

Highlights, comments, drawing tools, and other annotations can be stored as additional objects. Cleaning up the visible page does not automatically remove review notes.

If the PDF has passed through an internal review process, check the comments or annotations view before sharing it externally.

Form fields and stored values

Fillable PDFs can contain text fields, checkboxes, dropdowns, and other interactive elements. Values may remain stored even when the document looks like a normal PDF at first glance.

Reopen the saved file to confirm what it actually contains. Priviot also has a dedicated guide to filling PDF forms.

Embedded attachments

A PDF can carry additional files such as spreadsheets, images, or other documents. An embedded attachment is separate from the visible page sequence and can be easy to miss during a normal page-by-page review.

OCR and invisible text layers

A scanned document usually begins as images. Optical character recognition can add searchable text behind those images, making search and copy-and-paste possible.

That hidden text matters when redacting information. Covering the visible image does not necessarily remove an OCR text layer. A proper PDF redaction process needs to account for the document structure rather than just the visible color on the page.

Bookmarks, links, and document structure

PDFs can contain bookmarks, internal destinations, and external links. These are usually harmless, but reused templates can sometimes carry old filenames, internal URLs, or outdated destinations.

Fonts and embedded resources

PDFs may embed fonts or parts of fonts so the document renders consistently across devices. Images and other resources are also stored in the file. This is not usually a privacy issue, but it explains why two visually similar PDFs can have very different file sizes and internal structures.

What should you check before sharing a sensitive PDF?

  • Review document properties and relevant metadata.
  • Check comments and annotations.
  • Confirm stored values in forms.
  • Look for embedded attachments.
  • Check whether scans contain OCR text.
  • Review links and bookmarks.
  • After redaction, search for removed terms and try copying text from the redacted area.

Local editing does not automatically remove metadata

Editing a PDF locally avoids sending the file to an online editor for that operation. It does not automatically strip metadata, comments, or attachments. Local processing and document sanitization are separate questions.

Priviot PDF performs its core PDF work locally on your device. Priviot's technical transparency documentation explains its data flows and limitations in more detail.

Conclusion

A PDF is a container, not just a set of visible pages. For ordinary reading or printing, the distinction may not matter. For confidential, anonymized, or public documents, it is worth checking metadata, comments, forms, attachments, and hidden text layers as well.