Skip to content

About

Explainable, signature-aware PDF tampering and provenance analysis (pikepdf + pyHanko). Recovers earlier revisions, detects changes after signing, hidden content and active content.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

pdf-forensics-kit

CI Python License

Find out whether a PDF or Office document has been changed, what was changed, and whether it hides anything. Every conclusion is explained and backed by evidence, with no black-box score.

pdfforensics is a local command-line tool for fraud investigators, compliance and AML analysts, auditors, and anyone who receives invoices, statements, contracts or certificates and needs to decide whether to trust them. Point it at a file or a folder and you get:

  • a verdict (from no-indicators to strong-indicators),
  • a plain-language summary of what was found, what was ruled out and what to do next,
  • evidence you can show others, such as the text removed and added between an earlier version and the current one,
  • the earlier versions themselves, recovered as separate PDFs when the file still contains them.

Everything runs offline on your machine. The document is never uploaded, modified or executed.

What it is not: structural analysis cannot prove that the content of a document is true, and a clean result is not proof of authenticity. The tool tells you where to look and how strong the signal is. For a final answer, see How to read the results.


Contents


See it in action

pdfforensics interactive shell

A case folder contains three documents: a signed invoice, a Word offer and a bank statement.

pdfforensics analyze ./case-123 --out-dir ./case-123-reports

batch-summary.md (excerpt):

3 document(s) submitted: 2 need attention, 1 without significant findings.

Needs attention (most severe first)

  • invoice.pdf: strong-indicators - Revision 3 changes page content after the document was signed
  • offer.docx: review-recommended - Tracked changes still stored in the document (1 insertion(s), 1 deletion(s))

invoice.pdf.4f716821876d.summary.md (excerpt):

invoice.pdf: strong-indicators - 3 finding(s) need attention (1 critical, 1 high, 1 medium).

1-page PDF 1.3 produced by Acme Billing 4.2; pyHanko 0.37.0; 2 later edit(s) appended to the file; 1 signature(s), 1 with the signed bytes verified intact, but data was added after the last signature.

Key findings

  • [CRITICAL] Revision 3 changes page content after the document was signed: removed: "Total: 100 SEK"; added: "Total: 9 SEK"
  • [HIGH] Changes after signature 'Sig1' go beyond what the signer allowed: pyHanko: INTACT:UNTRUSTED,EXTENDED_WITH_OTHER,ILLEGAL_MODIFICATIONS
  • [MEDIUM] 324 byte(s) were appended after the last signature: signed up to byte 6,244 of 6,568

Checked and found in order

  • No hidden text, print-only annotations or hidden layers.
  • No JavaScript, launch actions, form submission or suspicious attachments.

Recommended next steps

  1. Extract the earlier revision(s) with pdfforensics extract-revisions and compare them with the current version. Ask the issuer for the original document.
  2. …

offer.docx: the Word document looks clean on screen, but its tracked changes still hold the original amount:

  • [MEDIUM] Tracked changes still stored in the document: deleted: "100 SEK"; inserted: "900 SEK"; by Mallory

Recover the version that was signed, and prove the difference:

$ pdfforensics extract-revisions invoice.pdf --out-dir revs
revs/invoice.rev01-of-03.pdf  (896 bytes)
revs/invoice.rev02-of-03.pdf  (6,244 bytes)     <- the signed version
revs/invoice.rev03-of-03.pdf  (6,568 bytes)     <- what you received

$ pdfforensics compare revs/invoice.rev02-of-03.pdf invoice.pdf
- page 1 only in A: `Total: 100 SEK`
- page 1 only in B: `Total: 9 SEK`

(The example documents are synthetic. They are generated by the test suite.)

Try it on a realistic case

examples/bank-statement/ contains a fictitious bank statement (watermarked SPECIMEN) and five forged or control versions: edited in a desktop editor, edited with an online editor, signed and then edited, and an Excel export with the original hidden in a very-hidden sheet. It also documents what the tool finds in each one:

pdfforensics analyze examples/bank-statement --summary

What it checks

PDF

Question How it is answered
Was the file edited after it was created, and what changed? Walks the PDF's internal revision chain. Every earlier version stored inside the file is rebuilt, and the tool lists which objects changed (page content, annotations, form fields, metadata) and diffs the visible text between versions.
Was it changed after it was signed? Is the signature genuine? Maps each signature's signed byte range to the revisions. It flags content changed after signing, validates the signature cryptographically with pyHanko, and optionally cross-checks with Poppler pdfsig (integrity, certificate trust and revocation reported separately). A signature on its own is not treated as tampering.
Does it show different things to different readers? Invisible text (with a check that tells OCR'd scans apart from hidden text), annotations that appear only when printed, hidden annotations, layers that are off by default or differ between screen and print.
Do the metadata make sense? Modification before creation, dates in the future, Info-dictionary vs XMP disagreements (timezone-aware), producer changes, XMP edit history, and consumer or online PDF editors (iLovePDF, Smallpdf, Sejda, …) in the tool chain.
Is the file structure sound? Data before the header or after the end of the file, repairs needed to open it, leftover unreferenced content, broken revision chains, encryption.
Is it dangerous to open? JavaScript, launch actions, form submission or data import, embedded files (e-invoice XML recognised as benign), XFA forms, multimedia, external links.
Where does it come from? A production fingerprint (software, PDF version, internal structure, fonts). Documents that claim the same issuer but were built differently stand out in a batch.

Word, Excel and PowerPoint (DOCX / XLSX / PPTX, including macro-enabled variants)

Question How it is answered
Was the wording or an amount changed? Unaccepted tracked changes, including the deleted text still stored in the file, the inserted text and the author names.
Is something hidden? Hidden Word text, hidden and very-hidden Excel sheets (very-hidden sheets cannot be unhidden from the Excel interface), hidden slides, comments.
Do the document properties add up? Modified before created, printed before created, dates in the future, author vs last editor, many saves with zero editing time.
Is it dangerous to open? Macros (VBA / Excel 4), with an extra flag when a macro-free .docx/.xlsx contains them. Also DDE fields, remote-template injection, content loaded from outside the file, embedded and ActiveX objects.
Is the file what it claims to be? Extension vs actual content, a ZIP disguised as an Office file, duplicate internal parts, path-traversal entries, decompression bombs, inconsistent package manifests.

For all formats, the file extension is checked against the real content (for example a PDF renamed to .docx).


What you get

Output files

With --out-dir, each document produces three files, named after the file and the first 12 characters of its SHA-256. Same-named files from different folders therefore never overwrite each other:

File For whom Content
invoice.pdf.forensics-report.md (next to the document) Everyone Written when you answer yes after an analysis, or with --save-report: the full report in Markdown, in the document's own folder
invoice.pdf.4f716821876d.summary.md Case handlers, managers One-page plain-language summary: headline, document profile, key findings with their evidence, what was ruled out, next steps, limitations
invoice.pdf.4f716821876d.forensics.md Reviewers Full report: file identity and hash, tool versions, summary, then every finding with explanation, evidence and benign explanations, and the verdict rule
invoice.pdf.4f716821876d.forensics.json Systems, archiving, re-analysis Everything, machine-readable (see below)
batch-summary.md Everyone (when several files are analysed) Executive summary, documents needing attention first, a table of all documents, and grouping by production pipeline

Without --out-dir, the Markdown report is printed to the terminal. --summary prints only the summary, and --json - prints JSON.

The verdict

Level Label Meaning
info no-indicators Nothing found in the checks performed
low minor-anomalies Anomalies that are usually benign; review them if the document matters
medium review-recommended Needs an explanation before you rely on the document
high significant-indicators Post-creation modification or hidden content
critical strong-indicators For example, content changed after signing or a broken signature

The rule is simple and printed in every report. Each finding has a severity and a confidence. A low-confidence finding counts one level lower, and the verdict is the highest level that remains. If any check failed or a format is only partly supported, the verdict is marked incomplete. A failed check never counts as "clean".

Anatomy of a finding

Every finding answers what, how sure and what else could explain it:

{
  "id": "revisions.content-changed",
  "title": "Revision 3 changes page content after the document was signed",
  "severity": "critical",
  "confidence": "high",
  "category": "revisions",
  "explanation": "An incremental update replaced or added page content streams ... The earlier version is still inside the file and can be extracted with 'pdfforensics extract-revisions'.",
  "evidence": {
    "revision": 3,
    "objects_changed": 1,
    "changed_roles": { "content": 1 },
    "text_removed": [ { "page": 1, "text": "Total: 100 SEK" } ],
    "text_added":   [ { "page": 1, "text": "Total: 9 SEK" } ]
  },
  "benign_explanations": [
    "Legitimate re-save by an editor after corrections",
    "Page added by an approved workflow (e.g. appended cover sheet)"
  ]
}

The JSON report

Key Content
file Name, absolute path, format (pdf / ooxml / ole), kind (pdf/docx/xlsx/pptx), size, SHA-256, MD5, analysis time (UTC)
tool Tool version and the versions of pikepdf, qpdf, pypdf, pyHanko and pdfsig, plus platform and whether external tools and network access were enabled
verdict Level, label, meaning, completeness, counts per severity, top findings, the rule, the disclaimer
summary The plain-language summary as structured data (headline, document profile, key findings, ruled out, recommended actions, limitations)
findings All findings, most severe first
facts Raw observations per analyser: revisions and per-revision diffs, signatures, metadata, page statistics, active content, Office properties, fingerprint
errors Checks that could not run, and why

Saved JSON can be re-summarised at any time: pdfforensics summarize *.forensics.json. It reproduces the summary exactly.

Exit codes

0 success · 1 a verdict reached the --fail-on level, or (with --fail-on-incomplete) an analysis was incomplete or a file could not be analysed · 2 input error (not found, unsupported, unreadable)


Install

Requires Python 3.10 or later.

git clone https://github.com/overjoyde/pdf-forensics-kit.git
cd pdf-forensics-kit
python3 -m venv .venv
.venv/bin/pip install -e '.[signatures]'      # add ,report for PDF reports, ,dev for the test suite
.venv/bin/pdfforensics --version
Component Required? Purpose Licence
pikepdf (qpdf) yes Low-level PDF structure MPL-2.0
pypdf yes Text of each PDF revision BSD
pyhanko optional ([signatures]) Cryptographic PDF signature validation MIT
reportlab optional ([report]) PDF report output BSD
Poppler pdfsig optional (brew install poppler / apt install poppler-utils) Second signature validator: trust and revocation GPL (separate program, called as a tool)

Office analysis uses only the Python standard library. There is no AGPL code and nothing phones home.


Usage

# Interactive shell (see below)
pdfforensics

# One file: full Markdown report in the terminal
pdfforensics analyze invoice.pdf

# Save the report as Markdown next to the document (asked interactively, or without asking:)
pdfforensics analyze ~/Case/invoice.pdf --summary --save-report     # -> ~/Case/invoice.pdf.forensics-report.md

# Just the plain-language summary (terminal, or write it to a file)
pdfforensics analyze invoice.pdf --summary
pdfforensics analyze offer.docx --summary offer-summary.md

# A whole case folder (recursively): per-file JSON + report + summary, plus a batch summary
pdfforensics analyze ./case-123 -r --out-dir ./case-123-reports

# Recover every earlier version of a PDF as a standalone file
pdfforensics extract-revisions invoice.pdf --out-dir ./revisions

# Compare a questioned copy with a reference copy (identity, metadata, production, text)
pdfforensics compare reference.pdf questioned.pdf

# Rebuild summaries later from saved JSON reports
pdfforensics summarize ./case-123-reports/*.forensics.json -o case-123-summary.md

# Automation: JSON on stdout, exit code 1 if anything is 'high' or worse
pdfforensics analyze ./inbox --json - --fail-on high
Option Effect
-r, --recursive Include sub-folders
--out-dir DIR Write JSON, report and summary per file (plus batch-summary.md)
--json FILE / --markdown FILE Write the JSON or Markdown report (- for stdout)
--summary [FILE] Output only the plain-language summary
--pdf-report PATH Write a summarised PDF report with the verdict, key findings and the edit timeline (a directory when several files are analysed). Needs pip install 'pdf-forensics-kit[report]'
--hide-info Leave info-level findings out of the Markdown
--save-report Save <document>.forensics-report.md next to each document without asking
--no-prompt Never ask whether to save a report (for scripts). In a terminal, analyze otherwise asks after the analysis
--fail-on LEVEL Exit with 1 if any verdict is at or above low/medium/high/critical
--fail-on-incomplete Exit with 1 if any analysis is incomplete (a check failed) or a file could not be analysed. Legacy .doc/.xls/.ppt files are always incomplete
--password PW User password for encrypted PDFs
--no-external-tools Do not run Poppler pdfsig, even if installed
--online-revocation Let pdfsig contact OCSP servers to check revocation (network access, off by default)
--max-pages, --max-objects, --max-revisions Scan limits for very large files. Hitting a limit is reported, never hidden

The analysis is also available from Python: from pdfforensics import analyze_file; analyze_file("x.pdf").to_dict().


Interactive shell

Run pdfforensics without arguments (or pdfforensics shell) to open a session that stays live for input:

The pdfforensics interactive shell analysing a forged bank statement

(Screenshot: the shell analysing the fictitious 05_signed_then_edited.pdf. The middle of the summary is cropped.)

The banner shows a gradient logo, the version, and which validators are active (pyHanko, pdfsig, offline or online, where reports are saved). Verdicts and severities appear as coloured chips and long lines wrap to the terminal width. The display adapts to the terminal: true colour when COLORTERM=truecolor, otherwise the 256-colour palette. It falls back to plain ASCII if the terminal can't show Unicode, and to no colour with NO_COLOR=1 or when output is piped.

Drag and drop a file or folder into the terminal and press Enter. It is analysed right away, and paths with spaces work.

Save the report when the analysis is done. After every analysis the shell asks:

? Save the report as Markdown in /Users/you/Case 42? [y/N] y
✓ report saved: /Users/you/Case 42/invoice.pdf.forensics-report.md

The report is always a Markdown file, saved in the same folder as the document, named <document>.forensics-report.md. It contains the full report: verdict, file hash, summary and every finding with its evidence. Enter or n saves nothing, and save saves the last analysis later. An existing report is never overwritten (a timestamp is added instead), and the document itself is never modified. For a folder, each document gets its report next to it, plus a forensics-batch-report-<time>.md in the common folder. set save always saves without asking, and set save never stops the question. Tab completes commands and file paths, the arrow keys recall earlier commands (history is kept in ~/.pdfforensics_history), and Ctrl-C interrupts a running analysis without leaving the shell.

Command What it does
<path> Analyse a dragged or typed file or folder and show the summary
analyze PATH... (a) Same, for several paths
report PATH Full report: every finding with its evidence
last [full] Show the previous result again
findings [LEVEL] Compact list of the previous analysis' findings at or above info/low/medium/high/critical
extract PDF [DIR] Recover every earlier revision of a PDF
compare A B Compare two PDFs
summarize JSON... Rebuild summaries from saved reports
save Save the last report(s) as Markdown next to the document(s)
set save ask|always|never After each analysis: ask to save the report (default), always save it, or never
set outdir DIR|off Also save JSON, report and summary files for every analysis in DIR
set external on|off Use Poppler pdfsig as a second signature validator (default on)
set online on|off Allow pdfsig to contact OCSP servers (network access, default off)
set recursive on|off Include sub-folders (default on)
status Session settings and available validators
help [COMMAND], clear, exit / quit / Ctrl-D

The shell also reads commands from a pipe, e.g. printf 'analyze a.pdf\nexit\n' | pdfforensics.


How to read the results

Commonly benign, so do not treat these as red flags on their own:

  • CreationDate differs from ModDate.
  • Incremental updates that only add a signature, form values or long-term-validation data.
  • Invisible OCR text over a scanned image.
  • Author differs from last editor.
  • Tracked changes in a document that is still under review.
  • Hidden helper sheets in spreadsheet templates.
  • Unused fonts left behind by generators.

High-signal findings, which deserve follow-up:

  • revisions.content-changed with text evidence: the earlier version is recoverable.
  • Content changed after signing, signature.broken, pdfsig.integrity-failure.
  • office.tracked-changes with deleted text: the original wording or amount is still in the file.
  • metadata.editor-tool on a document that supposedly came straight from a bank, insurer or ERP system.
  • signature.validators-disagree: the two validators give different answers, so get manual expert review.

Settling the question. Structural analysis is the weakest of the usual sources of evidence. When a document matters, go up this list (the summary suggests it automatically):

  1. A valid digital signature from a trusted certificate, covering the relevant version
  2. A known-good hash from the issuing system or an immutable archive
  3. The source system's audit log or version history
  4. An independently obtained copy from the issuer
  5. This tool's structural findings

Never report no-indicators as "authentic" or "not tampered". It means the checks found nothing.


Safety and privacy

  • The document is never modified. The only file ever written next to it is the Markdown report you explicitly confirm (or request with --save-report / set save always), and existing reports are never overwritten.
  • Read-only, captured once. The file is read once without following symbolic links, then hashed and analysed from those bytes. A file that changes while it is read is rejected, and a file that changes during analysis is reported.
  • Offline. No uploads and no telemetry. pdfsig runs with -no-ocsp unless you pass --online-revocation.
  • Nothing is executed or fetched. Macros, JavaScript, DDE fields, embedded files and external links are only inspected.
  • Hostile-file hardening. Size limits, bounded ZIP/XML parsing, bounded PDF stream decoding (XMP 2 MB, page content 64 MB per page and 256 MB per document, other streams 64 MB; reported as *.stream-too-large / metadata.xmp-too-large), XML entity declarations refused, and external tools given a private temporary copy with no shell.
  • Reports are sensitive. They quote metadata and changed text, so give them the same classification and handling as the documents themselves.

Supported formats and limitations

Format Support
PDF (all versions, classic/stream/hybrid xref, linearized, encrypted with password) Full
DOCX, DOCM, DOTX, DOTM Full (Word checks)
XLSX, XLSM, XLTX, XLTM Full (Excel checks)
PPTX, PPTM, POTX, PPSX … Full (PowerPoint checks)
Legacy .doc / .xls / .ppt (OLE2) Recognised only; the report is marked incomplete (use oletools)
Images, e-mail, other ZIPs Not supported (rejected with a clear message)

Known limits:

  • Office XML signatures are detected but not validated cryptographically.
  • There is no image forensics (e.g. error-level analysis) and no font-glyph analysis.
  • Signature trust depends on the certificates configured for pyHanko or Poppler.
  • It cannot judge whether the content is true.

Finding reference

All finding ids by analyser (click to expand)
Analyser Findings
input (all formats) extension-mismatch, changed-after-capture
PDF structure data-before-header, data-after-eof, extra-eof-markers, repaired, unreachable-objects, encrypted, xref-chain-broken, unlinked-revision
PDF revisions content-changed (with text diff), annotation-or-form-update, metadata-update, signature-update, other-update, not-all-compared
Timeline timeline.inconsistent-times
PDF signatures (pyHanko) intact, broken, disallowed-modification, bytes-after-last-signature, malformed-byterange, usage-rights, not-validated, validation-error, unparseable
PDF signatures (pdfsig, optional) integrity-ok, integrity-failure, integrity-unknown, certificate-revoked, certificate-expired; signature.validators-disagree
PDF metadata modified-before-created, future-date, info-xmp-date-mismatch, producer-mismatch, editor-tool, manipulation-library, xmp-history, absent
PDF content invisible-text, ocr-text-layer, print-only-annotations, hidden-annotations, layer-view-print-differs, layers-hidden-by-default
PDF active content javascript, launch-action, submit-or-import, remote-goto, multimedia, xfa, embedded-files, e-invoice-attachment, additional-actions, uris
Office container not-an-office-package, extension-mismatch, duplicate-parts, unsafe-paths, encrypted-parts, compression-bomb, resource-limit, opc-nonconformant, parts-not-parsed, legacy-format
Office metadata modified-before-created, printed-before-created, future-date, different-editor, zero-edit-time, remote-template-name, metadata-absent
Office content tracked-changes, track-changes-enabled, hidden-text, very-hidden-sheets, hidden-sheets, external-workbook-links, hidden-slides, comments
Office active content macros, remote-template, dde-field, external-content, embedded-objects, activex, hyperlinks, xml-signature
fingerprint Facts only: producer family, pipeline hash, fonts, filters, xref style / application, version, company, template

Documentation, development and credits

  • docs/METHODOLOGY.md: evidence handling, the findings model, what is normal, stronger verification sources.
  • docs/VALIDATION.md: results on public PDF and Office samples, and the defects found and fixed along the way.
  • docs/ANALYSIS.md: review of the upstream project and why this is a rewrite.
  • Agent skill: .agents/skills/document-forensics/ teaches coding agents (Codex, opencode, …) to run the tool safely and report without over-interpreting.

Tests. .venv/bin/pip install -e '.[signatures,dev]' && .venv/bin/python -m pytest. All fixtures are generated synthetically at test time (tests/pdfgen.py, tests/officegen.py), including a signed PDF made with a throw-away certificate. The repository contains no real documents. CI runs on Linux, macOS and Windows, with and without pyHanko, and with Poppler on Linux.

Credits. Inspired by Rlahuerta/pdf-forensics-toolkit (MIT). The code was written anew; see NOTICE. Released under the MIT License.

About

Explainable, signature-aware PDF tampering and provenance analysis (pikepdf + pyHanko). Recovers earlier revisions, detects changes after signing, hidden content and active content.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages