Document Forensics & Cybersecurity•Published October 2, 2026•17 min read

PDF Metadata Forensics: Inspecting Author, Dates & Invisible Revision Streams

Every business day, millions of PDF files containing legal contracts, financial audits, medical case files, executive memoranda, and academic research papers are emailed, uploaded to public repositories, and filed in court proceedings. Most users assume that what they see on screen is the entirety of the document's contents. In reality, Portable Document Format (PDF) files frequently contain extensive invisible metadata, historic document versions, author identities, and software footprints that pose severe privacy and legal liabilities.

High-profile leaks in investigative journalism, government disclosures, and legal discovery frequently trace back to hidden PDF metadata. In several notorious cases, documents released with black-box redactions over sensitive text were instantly decrypted because the redaction tool merely drew a black rectangle over the text rather than scrubbing the underlying content stream, or because an unredacted draft was preserved inside an Incremental Revision Stream.

In this technical masterclass, we conduct a deep-dive forensic examination of the PDF binary architecture. We explore how metadata is dual-stored in both the Document Information Dictionary and the XMP Metadata Stream, analyze how incremental saves preserve deleted objects, and demonstrate how to inspect and sanitize PDFs client-side without transmitting confidential files to external servers.

Advertisement
Responsive In-Article Ad Slot

The Dual Metadata Architecture: Info Dict vs. XMP Streams

The PDF specification (ISO 32000-2:2020) defines two distinct subsystems for storing document-level metadata:

PDF METADATA DUAL-STORAGE ARCHITECTURE 1. Classic Info Dictionary (PDF 1.0+) Stored in PDF Trailer Dictionary /Title (Project Titan Roadmap) /Author (Jane Doe <jdoe@corp>) /Creator (Microsoft Word 365) /Producer (macOS Quartz PDF) /CreationDate (D:20261002...) 2. Adobe XMP Packet (PDF 1.4+ / ISO) XML Stream in Document Root Catalog <?xpacket begin="" id="W5M0Mp..."?> <x:xmpmeta xmlns:x="adobe:ns:meta/"> <rdf:RDF><rdf:Description> <dc:creator><rdf:Seq>... <xmpMM:DocumentID>uuid:4a9...
Figure 1: Comparison between the legacy trailer Info Dictionary and the XML-based XMP metadata packet stored in the Document Catalog.

1. Document Information Dictionary

The Info Dictionary is referenced by the /Info key within the document trailer. It consists of primitive PDF string objects:

  • /Title and /Subject: The semantic title and abstract.
  • /Author: The user profile or OS login name of the file creator.
  • /Creator: The original authoring application (e.g., Adobe InDesign 19.0, Microsoft Word).
  • /Producer: The PDF conversion engine (e.g., Adobe PDF Library, Ghostscript, Quartz).
  • /CreationDate and /ModDate: Exact timestamps formatted as (D:YYYYMMDDHHmmSSOHH'mm'), where O indicates the UTC offset (+, -, or Z).

2. Extensible Metadata Platform (XMP) Packet

Introduced by Adobe and standardized under ISO 16684-1, XMP stores structured metadata as an RDF/XML data stream inside the PDF Catalog dictionary (/Type /Metadata /Subtype /XML). XMP supports standardized schemas including Dublin Core (dc:), Adobe PDF Schema (pdf:), and XMP Media Management (xmpMM:).

Crucially, the xmpMM:History and xmpMM:InstanceID tags can contain an entire audit log of every software package, machine GUID, and prior document version that contributed to the current file.

Inspect PDF Metadata & Hidden Headers

Extract all Info dictionary keys, creation timestamps, producer tools, and raw XMP packets instantly in your browser.

Launch In-Browser PDF Metadata Viewer →

Forensic Threat: PDF Incremental Updates & Revision Streams

Perhaps the most insidious data leak vulnerability in PDF architecture is the Incremental Update mechanism. In traditional desktop PDF editors, when a user edits text, adds comments, or deletes a confidential page, the software frequently avoids rewriting the entire file from scratch to improve save speed.

PDF INCREMENTAL UPDATE BYTE STREAM TIMELINE Revision 1 (Original Draft) Obj 1: Catalog Obj 2: Unredacted Contract Obj 3: Confidential PII xref table 1 trailer → startxref 1024 Persists in Bytes 0 - 24KB Revision 2 (Redacted Save) Obj 4: Black Box Annotation Obj 5: Updated Page Ref xref table 2 (New pointers) trailer → /Prev 1024 startxref 4096 Appended to Bytes 24KB - 40KB Forensic Extraction Parsing earlier XREF table instantly recovers Obj 3 "Confidential PII" Zero cracking required Requires Linearized Sanitization
Figure 2: Incremental updates append new objects and trailers at the end of the file, leaving previous drafts and unredacted PII intact in earlier byte ranges.

How Revision Trees Work in Binary Streams

When an incremental save occurs, the editor writes a new body block, an updated cross-reference (xref) table, and a new trailer dictionary containing a /Prev key pointing to the byte offset of the previous trailer. Standard PDF viewers only read the latest trailer and render the most recent state. However, any forensic tool or hex editor can navigate backwards along the /Prev chain to recover every historical iteration of the document.

JavaScript Code: In-Browser PDF Metadata Extractor

Using client-side JavaScript, you can parse the raw binary ArrayBuffer of any PDF file to extract the Information Dictionary and raw XMP stream directly in the browser:

/**
 * Extracts raw Info dictionary and XMP metadata from a PDF ArrayBuffer
 * @param {ArrayBuffer} buffer - The PDF binary buffer
 * @returns {Object} Extracted metadata properties
 */
export function inspectPdfMetadata(buffer) {
  const decoder = new TextDecoder('latin1');
  const text = decoder.decode(buffer);

  const results = {
    pdfVersion: null,
    hasIncrementalUpdates: false,
    infoDict: {},
    xmpPacket: null
  };

  // 1. Extract PDF Version Header
  const headerMatch = text.match(/%PDF-(\d+\.\d+)/);
  if (headerMatch) results.pdfVersion = headerMatch[1];

  // 2. Detect multiple trailers (Incremental Updates)
  const trailerMatches = text.match(/trailer\s*<<[\s\S]*?>>/g) || [];
  results.hasIncrementalUpdates = trailerMatches.length > 1;

  // 3. Extract Info Dictionary
  const infoMatch = text.match(/\/Info\s+(\d+)\s+(\d+)\s+R/);
  if (infoMatch) {
    const objNum = infoMatch[1];
    const genNum = infoMatch[2];
    const objRegex = new RegExp(`${objNum}\\s+${genNum}\\s+obj\\s*<<([\\s\\S]*?)>>`, 'm');
    const objBody = text.match(objRegex);
    if (objBody && objBody[1]) {
      const keys = ['Title', 'Author', 'Subject', 'Keywords', 'Creator', 'Producer', 'CreationDate', 'ModDate'];
      keys.forEach(k => {
        const valMatch = objBody[1].match(new RegExp(`/${k}\\s*\\((.*?)\\)`, 's'));
        if (valMatch) results.infoDict[k] = valMatch[1];
      });
    }
  }

  // 4. Extract Raw XMP Packet
  const xmpStart = text.indexOf('<?xpacket begin');
  const xmpEnd = text.indexOf('<?xpacket end', xmpStart);
  if (xmpStart !== -1 && xmpEnd !== -1) {
    results.xmpPacket = text.substring(xmpStart, xmpEnd + 19);
  }

  return results;
}

Best Practices for PDF Document Sanitization

Before distributing or publishing sensitive PDF documents, organizations should enforce the following security checklist:

  1. Destructive Redaction Scrubbing: Ensure redaction tools permanently erase the underlying text characters, vector outlines, and bitmap pixel regions from the PDF content stream rather than simply overlaying black rectangular vector shapes.
  2. Purge Both Info and XMP: Sanitization utilities must clear both the trailer /Info dictionary and the Catalog /Metadata XMP packet simultaneously to prevent metadata discrepancies.
  3. Full Document Linearization (Save As / Clean Export): Always perform a complete document re-serialization (such as PDF "Linearize" or full re-write) rather than an incremental save. This physically deletes all historic revision streams, detached unreferenced objects, and /Prev trailer pointers.
  4. Scrub Embedded Attachments and Annotations: Check for and remove invisible embedded file attachments (/EmbeddedFiles), comments, sticky notes, and hidden print-only layers.

Frequently Asked Questions

Yes. Even if /Author is blank, the embedded XMP stream may contain dc:creator, xmp:ModifyDate, xmpMM:History, or unique machine GUIDs (xmpMM:DocumentID). Additionally, embedded font subsets, printer calibration metadata, or operating system file paths in embedded images can uniquely identify the originating workstation.
Yes, rendering PDF pages to high-resolution raster images (e.g., PNG) and re-assembling them into a fresh PDF completely destroys all underlying object streams, XMP packets, revision histories, and hidden text layers. However, this also removes searchable OCR text and increases file size unless OCR is re-applied.
/Creator refers to the original software application in which the document content was authored (e.g., Microsoft Word, Google Docs, Apple Pages), whereas /Producer refers to the PDF conversion engine that synthesized the PDF binary file (e.g., Adobe Acrobat Distiller, Skia / PDF, wkhtmltopdf).
If opened in desktop software like Adobe Acrobat, malicious PDFs can contain embedded URIs, Form Submission actions (/SubmitForm), or JavaScript that phones home to a remote server. However, analyzing the PDF with client-side WebAssembly tools or sandboxed text parsers prevents any external network callbacks.

Summary & Key Takeaways

Document metadata is a double-edged sword: invaluable for indexing and archiving, but hazardous for document confidentiality and legal compliance. By integrating client-side metadata inspection and automated document sanitization workflows into your organizational publishing pipelines, you can protect sensitive information and eliminate forensic liabilities before sharing documents publicly.

CS

Collabsource Editorial Team

Security researchers and document engineering specialists dedicated to client-side data privacy, binary file analysis, and digital forensics.