Dev.to Security 🔐 Cybersecurity 👁 0 📖 9 min read

Freight Packet PDF Contains Redacted Text (When Signatures Must Survive)

Short answer: if a redacted PDF still contains extractable text, it has an overlay rather than a completed redaction; remove the underlying content, inspect the final bytes independently, reassemble the authorized freigh

Short answer: if a redacted PDF still contains extractable text, it has an overlay rather than a completed redaction; remove the underlying content, inspect the final bytes independently, reassemble the authorized freight pages, and sign only that finished artifact.

Treat a redaction as complete only when the sensitive content has been removed from the PDF's reachable objects and the released bytes pass an independent extraction check. A black rectangle is evidence of appearance, not removal. In a freight packet pipeline, split first, identify the pages and fields under policy, redact by rewriting content, sanitize related objects, reassemble, and then sign the final release artifact. Preserve the signed original separately; changing any byte in a signed PDF can invalidate its signature.

The decision rule is strict: if the old value can be recovered by text extraction, copy and paste, object inspection, layer changes, annotation removal, or decoding an embedded asset, the document is not redacted. Do not ship it.

Why does a redacted PDF still contain text after an overlay?

PDF describes a document through objects: page content streams, fonts, images, annotations, form fields, optional-content groups, metadata, attachments, and cross-reference data. What a viewer paints is only one interpretation of those objects. Drawing an opaque rectangle above a consignee name changes the visible composition, but the text-showing operation beneath it can remain in the page content stream. Selection and extraction tools may still read the name.

The same split between appearance and content occurs in scanned paperwork. Covering pixels on a rendered page does not prove that the original image, an OCR text layer, or an alternate image remains unreachable. A form field may look blank while its value survives in the field dictionary. A comment can cover text without deleting either item. Incremental PDF updates deserve equal suspicion because earlier object revisions may remain in the file even though a current viewer follows the newest cross-reference information.

This distinction matters during bundle work. A logistics packet might contain a bill of lading, a commercial invoice, delivery evidence, customs pages, and a carrier exception note. Different recipients are allowed to see different fields. Page 4 may expose a driver's phone number while page 11 contains a customer reference in both visible text and an AcroForm field. If the pipeline splits by page, paints boxes, and merges the pages again, it can produce a convincing preview and a failed disclosure control at the same time.

No screenshot settles this.

Stop there.

Build the release artifact, then test its bytes

Use an immutable source, a policy-driven transformation, and a distinct release artifact. The audit record should bind those stages without copying the secret into logs. A useful flow is: ingest and hash the source; verify any existing signature; split the logical packet; locate approved redaction targets; rewrite or rasterize the affected page according to policy; remove associated interactive and hidden material; merge the authorized pages; validate; sign the final bytes; then record the release hash and verification result.

Rasterization can reduce the number of PDF object types that must be inspected, but it is not a magic verb. The output must contain only the sanitized pixels, at a resolution that keeps required freight details legible, and it must not carry the source image, OCR layer, attachment, or editing history alongside the rendered result. Rebuilding structured content can preserve searchability and accessibility, but it requires more careful handling of text operators, clipping paths, forms, annotations, and resources. That is a real trade-off: a flatter artifact is easier to reason about, while a structured artifact retains more document behavior.

Here is a small TypeScript gate around generic transformation and inspection adapters. Matching a string is not a complete redaction proof. The example makes the useful minimum explicit: the final artifact is independently inspected, signature order is enforced, and audit events contain hashes and field identifiers rather than sensitive values.

import { createHash } from "node:crypto";

type Artifact = { bytes: Uint8Array; mediaType: "application/pdf" };
type Target = { page: number; fieldId: string; forbiddenText: string };
type Finding = { page: number; source: string; text: string };

interface PdfTransformer {
  redact(source: Artifact, targets: Target[]): Promise<Artifact>;
  merge(parts: Artifact[]): Promise<Artifact>;
  sign(source: Artifact): Promise<Artifact>;
}

interface PdfInspector {
  extractReachableText(source: Artifact): Promise<Finding[]>;
  listEmbeddedFiles(source: Artifact): Promise<string[]>;
  hasIncrementalHistory(source: Artifact): Promise<boolean>;
}

const sha256 = (bytes: Uint8Array) =>
  createHash("sha256").update(bytes).digest("hex");

async function produceRelease(
  source: Artifact,
  targets: Target[],
  transformer: PdfTransformer,
  inspector: PdfInspector,
  otherAuthorizedParts: Artifact[]
) {
  const redacted = await transformer.redact(source, targets);
  const merged = await transformer.merge([redacted, ...otherAuthorizedParts]);
  const findings = await inspector.extractReachableText(merged);
  const forbidden = new Set(targets.map((target) => target.forbiddenText));
  const leaks = findings.filter((finding) => forbidden.has(finding.text));
  const attachments = await inspector.listEmbeddedFiles(merged);
  const hasHistory = await inspector.hasIncrementalHistory(merged);

  if (leaks.length || attachments.length || hasHistory) {
    throw new Error("Release blocked by PDF disclosure checks");
  }

  const signed = await transformer.sign(merged);
  return {
    artifact: signed,
    audit: {
      sourceSha256: sha256(source.bytes),
      releaseSha256: sha256(signed.bytes),
      targetIds: targets.map((target) => target.fieldId),
      checkedPages: [...new Set(targets.map((target) => target.page))],
      result: "pass" as const
    }
  };
}

In production, avoid keeping forbiddenText longer than verification requires. A keyed digest can support exact-value checks when the extractor and policy service share a controlled normalization scheme, though normalization errors can create false confidence. Random shipment references also resist guessing better than phone numbers or common names. The test design must reflect the data.

Separate redaction proof from signature proof

A digital signature answers whether specified bytes changed after signing and connects the signature to a certificate and validation context. It does not certify that the signed content is safe to disclose. Redaction answers a different question: does the release artifact still expose prohibited information? Both checks are necessary, and their order matters.

If an inbound carrier document is already signed, retain that exact file and its verification evidence in the restricted archive. A disclosure copy derived from it is a new artifact with a new hash and provenance record. Redact and merge that copy, validate the result, and apply an outbound signature only after the byte sequence is final. Trying to preserve the inbound signature while editing covered content confuses chain of custody with byte integrity. The audit trail should link the source hash, transformation policy version, target identifiers, output hash, inspector version, signer identity, and timestamps. It should not imply that the source signer approved the transformed copy.

Merge order therefore stops being an implementation detail. Suppose a 23-page export packet is split into five logical documents. Two pages require redaction, one page carries the inbound signature, and the consignee receives only 14 pages. The release manifest should say which source pages produced those 14 pages and which policy acted on the two sensitive pages. If source page 4 becomes release page 2, and source page 11 becomes release page 7, the audit record must retain both mappings; a bare list of output page numbers cannot explain which input was transformed. The manifest also needs the transformation policy version and the hash of the exact source packet. Page numbering can change; stable manifest identifiers should not. The inbound signature remains evidence about the archived source, while the outbound signature covers the newly assembled 14-page release. Sign once, after assembly, then verify the signature against the stored release hash before delivery.

Order wins.

Never log the recovered secret to prove that the gate found it. Record a target ID, page, detection channel, policy version, and a protected digest or boolean result. Debug output often has broader retention and access than the document store, so a successful security check can otherwise create a second disclosure.

What should the verifier try to recover?

Start with multiple views of the final bytes. Extract text through a parser independent from the component that performed the rewrite. Inspect page content and form XObjects, annotations, AcroForm values, optional-content groups, embedded files, document metadata, and images. Check whether incremental revisions or unexpected trailing data remain. Render every released page as well, because a structurally clean file can still reveal text visually through a misplaced box, transparency, clipping, or a wrong page transform.

Search for exact forbidden values when policy permits, then add patterns for classes of data: telephone formats, account references, email addresses, and shipment identifiers. Exact matching alone misses changed whitespace, character encoding, ligatures, OCR substitutions, and text drawn one glyph at a time. Pattern matching alone produces noise. The practical gate combines target-aware assertions with structural inspection and visual review samples.

Test adversarial fixtures rather than one friendly invoice. Include rotated pages, cropped content, text inside reusable form objects, nested resources, transparent overlays, annotations, filled form fields, scanned pages with OCR, attachments, and files saved through incremental updates. Include a legitimate document whose ordinary prose resembles an identifier, too, or an overbroad detector will halt every batch. Each fixture needs a known expected result.

Channel Failure it catches Release evidence
Independent text extraction Covered or hidden text remains reachable Target IDs and pass/fail result
Object and attachment inspection Forms, annotations, revisions, or embedded sources survive Finding type and object reference
Full-page rendering Wrong coordinates, transparency, or visible remnants Render hash plus controlled review record
Signature validation Final bytes changed after approval Signature validation record and release hash
Manifest reconciliation Split or merge selected the wrong pages Source-to-release page mapping

False positives should block the release without destroying evidence. Quarantine the candidate artifact, keep the immutable source under its existing access policy, and route only identifiers to the operations queue. A retry is appropriate for a transient worker interruption; it is not an answer to a deterministic leakage finding.

Operate the pipeline as a disclosure control

Measure the system at document boundaries. Useful signals include input and output hashes, page counts before and after each authorized split or merge, target counts, inspection duration, render failures, signature validation status, and policy versions. Keep raw extracted text out of telemetry. A sudden change in average output size or in the ratio of inspected objects per page can be a regression signal, but it is not itself proof of leakage or safety.

Cost pressure is real for an independent builder processing large packets. Run cheap structural checks and targeted extraction on every artifact, while reserving expensive visual analysis for affected pages plus a policy-defined sample of the remainder. Do not sample the basic disclosure gate. Cache results only by exact artifact hash and verifier version; a page that looks identical may have different underlying objects.

Deployment needs a fixed corpus and a staged rollout. Before changing the PDF writer, renderer, OCR component, or merge logic, run the adversarial fixtures and compare both policy outcomes and rendered pages. During rollout, process candidates without releasing them until the old and new gates agree. Roll back on unexplained disagreement, because parser diversity is useful only when differences are investigated rather than averaged away.

The final operational checklist belongs in the release procedure: confirm the source hash and access class; verify inbound signatures before transformation; map source pages to recipients; apply the approved redaction policy; sanitize interactive, embedded, and historical content; merge only authorized pages; run independent structural, extraction, and rendering checks; sign the completed artifact; verify that signature; store the manifest and final hash; release only on a recorded pass.

A dark box can be drawn in milliseconds. The trustworthy result is the chain of evidence around the bytes: what entered, what policy changed, what independent checks could no longer recover, which pages were assembled, and exactly what was signed. For freight bundles, that chain protects both confidentiality and the audit trail without pretending one can substitute for the other.

References

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.