Skip to content

Measures, not assurances

Security, described specifically enough to be checked.

Every control on this page is implemented in the code that serves you this document. Where a number appears, it is imported from the module that enforces it rather than typed here, so the page cannot drift out of step with the software.

The unusual thing about securing this site is what is absent from the attack surface. There is no document-upload endpoint or server-side document storage. There is an anonymous aggregate analytics database and a restricted admin login with a server-only password hash and signing secret. File parsers, collector endpoints, credentials and served code each need appropriate protection.

Threat model

Local files and public endpoints need different defences.

For a conventional PDF service the worst realistic outcome is a data breach: somebody obtains the documents it holds. This site has no stored document copies to expose: no document bucket, processing queue or temporary file directory. That does not eliminate the risk of compromised browser code, administrator credentials or hosting infrastructure.

A document may misdeclare its type, carry a filename engineered for traversal, or contain content that exhausts memory. The input and engine checks reduce those risks. Separately, the collector validates bounded event fields, uses parameterized SQL and deduplicates delivery; browser reports cannot prove that a real PDF operation occurred.

Not in the surface
Document-upload endpoint, document object storage, processing queue or public-account database.
In the surface
Browser parsers, shipped code, the measurement collector and aggregate store, and admin authentication.

Response headers

What every response carries

These are set on responses from the origin, so they apply to the tool pages themselves rather than only to a documented ideal.

Content-Security-Policy
By default no third-party script origin is allowed. Optional GA or ad configuration adds named provider origins; their scripts separately require consent and non-admin eligibility. First-party measurement stays same-origin. CSP complements, rather than replaces, validation and secure dependencies.
Strict-Transport-Security
After a browser receives HSTS over HTTPS, it uses HTTPS for later requests during the policy lifetime. Protecting a first-ever visit additionally depends on HTTPS links and browser preload status; the header alone does not prove preload enrollment.
X-Content-Type-Options: nosniff
Stops the browser from second-guessing a declared content type. A response we serve as data cannot be re-interpreted as a script.
X-Frame-Options: DENY
The site cannot be embedded in a frame on another origin, which removes clickjacking — an overlay tricking you into clicking a tool control you cannot see.
Referrer-Policy
Restricts what is sent in the Referer header on outbound navigation, so the full path of the page you were on is not handed to third parties.
Permissions-Policy
Denies browser capabilities the tools do not need — camera, microphone, geolocation and similar — so a compromised script could not reach for them.
Cross-Origin-Opener-Policy
COOP severs the window relationship with cross-origin openers, isolating this tab’s browsing context from pages that opened or were opened by it.
Cross-Origin-Resource-Policy
CORP prevents other origins from embedding our resources as their own, limiting cross-origin read attacks against what we serve.

Checks around processing

Input validation, at the relevant stage

File-size and signature checks run before PDF parsing. Page counts require structural parsing, so page budgets are checked afterwards, before further processing. Validation reduces exposure but cannot prove that a document is harmless.

  1. 01

    Type confirmed by magic bytes

    A file extension and a browser-reported MIME type are unverified claims. We read the leading bytes and require the real signature — the %PDF- marker for documents, and the specific byte patterns for JPEG, PNG, WebP, BMP, GIF and TIFF. Something merely named .pdf is rejected.

  2. 02

    Empty and truncated input refused

    A zero-length file produces a clear error rather than an unhandled exception deep in a parser, which is both better behaviour and a smaller surface.

  3. 03

    Size budgets enforced

    Per-file and per-batch ceilings are applied to the selection before processing starts, so a run that cannot succeed fails immediately with an explanation instead of exhausting the tab.

  4. 04

    Page count checked

    After structural parsing, the page count is checked against the processing budget. An over-budget document is refused before page processing; this check does not prevent all parser-level memory exhaustion.

  5. 05

    Expansion ratio guarded

    For archives and compressed streams, an entry whose declared uncompressed size exceeds the per-file budget is refused, and one whose expansion ratio is far beyond legitimate content is skipped. This is the decompression-bomb defence.

  6. 06

    Filenames sanitised

    Output and archive-entry names are stripped to their final path component, with directory separators, traversal sequences, control characters, characters illegal on common filesystems and reserved Windows device names all removed, and total length capped while the extension is preserved. A name crafted to write outside its intended folder cannot.

Enforced limits

The actual numbers

Imported directly from the validation module that enforces them. If a limit changes in the code, this page changes with it — these cannot drift apart.

Per file

250 MB

A single input file above this is refused before a parser touches it.

Per batch

600 MB

The combined size of one run, since every file shares the tab’s memory.

Files per run

250

Stops a mis-selected directory from queueing thousands of documents.

Pages per document

5000

An absurd page count is refused rather than locking the tab rendering it.

These are memory limits rather than commercial ones. A server-based service accepts larger input because the memory being spent is its own; here the budget belongs to your browser tab, and exceeding it would crash the tab mid-task rather than merely slow it down.

In the codebase

Practices that keep the surface small

No eval of document content

Nothing extracted from a document is ever evaluated as code. Text, metadata, field values and embedded scripts are treated strictly as data, and the CSP would block the attempt even if the code tried.

Text is escaped, never injected

Document-derived strings are rendered through the framework’s escaping rather than inserted as markup, so a filename or PDF title containing a script tag displays as characters.

Engines loaded on demand

A tool imports only the engine it needs. The homepage ships no PDF parser and no WebAssembly, which keeps the code executing on any given page to the minimum that page requires.

Dependency auditing

Dependencies are kept deliberately few, audited for known advisories, and updated when a relevant one is published. Fewer packages is itself a security property.

Assets from our own origin

Core fonts, page scripts and styles are served from this origin. OCR loads its configured engine/model assets; optional GA and ad providers add separate dependencies only behind their gates. Asset origin policy and dependency integrity still need deployment verification.

Content-free reporting

Consented first-party events contain no filenames, options, arbitrary exceptions, IPs, user agents or visitor IDs. Minute aggregates last 200 days, grouped failure details 90 days and hashed random delivery receipts 24 hours; daily totals remain. Hosting logs need a separate retention policy.

Private administrator sessions

The versioned root-scoped cookie is HttpOnly, SameSite=Strict and Secure in production, with an eight-hour expiry. The collector independently excludes authenticated administrators in the same browser; private reporting endpoints authorize each request.

Failures are handled, not silent

Validation failures raise typed errors with an explanation and a suggested alternative. A tool that cannot safely proceed says so rather than producing a quietly wrong file.

Responsible disclosure

Found something? We would rather hear it from you.

hello@mypdfiles.com

Put security in the subject line so the report is triaged ahead of everything else. There is no bug bounty programme — this is a free site with no revenue to pay one from — but reports are taken seriously, acknowledged, and fixed.

  • Describe the steps to reproduce precisely, including the browser and its version.
  • Say what an attacker would actually gain, which helps us judge severity honestly.
  • Please allow a reasonable period for a fix before publishing the details.
  • Do not include a real confidential document — a synthetic file that triggers the same behaviour is more useful and carries no risk.
  • Please do not test by attacking the live site’s infrastructure or attempting denial of service.

If a fix is user-visible or changes a documented limit, the affected pages are updated as part of shipping it rather than afterwards.

Questions about security

What is the actual threat model here?

Local processing avoids server-side document storage, but does not remove every threat. A crafted file can target the parser in your tab; malicious clients can target the public measurement API; and the admin panel, dependencies and served code must also be protected. File validation, bounded telemetry validation and authenticated private reports address different parts of that surface.

Why validate by magic bytes rather than the file extension?

Because an extension and a browser-reported MIME type are both simply claims made by whoever named the file, and neither is verified by anything. We read the leading bytes and look for the real signature — %PDF- for a PDF, the specific byte patterns for JPEG, PNG, WebP, BMP, GIF and TIFF. A file called report.pdf that is not a PDF is rejected before a parser sees it.

What stops a decompression bomb?

Two independent checks. An entry whose declared uncompressed size exceeds the per-file budget is refused outright, and an entry whose expansion ratio is wildly beyond what legitimate content produces is skipped rather than expanded. A crafted archive that would otherwise balloon to fill available memory therefore fails as a handled error instead of crashing the tab.

Does the Content-Security-Policy really allow no third-party scripts?

By default, no third-party script origin is permitted. Optional GA or advertising configuration adds narrowly named provider origins; loading their scripts still requires consent and server-confirmed non-admin eligibility. The first-party reporting collector needs no external script origin. CSP is one defence, not a guarantee that shipped code or permitted scripts cannot be compromised.

Is my document scanned for malware?

No. The site performs format and resource-budget checks, not a comprehensive malware scan. The document is processed locally, so your browser’s parser encounters it. Validation and size limits reduce exposure but cannot prove a document harmless; only open files from sources you trust.

How do I report a vulnerability?

Email hello@mypdfiles.com with “security” in the subject line. Include the steps to reproduce, the browser and version, and what an attacker would gain. Please allow a reasonable period for a fix before publishing, and please do not include a real confidential document in your report — a synthetic file that triggers the same behaviour is more useful and carries no risk.