Skip to content

Quality · 7 min read · Updated 2026-09-19

Why is my PDF so large?

A diagnostic walkthrough rather than a list of tips: how to work out which part of a specific file is consuming the bytes before you change anything.

An oversized PDF is a mystery with a finite number of suspects, and guessing at them wastes effort and quality. Before you compress anything, it is worth spending two minutes finding out where the weight actually sits — because the fix for a file bloated by revision history is nothing like the fix for a file bloated by a full CJK font, and applying the wrong one degrades your document for no gain.

Why is my PDF so large?

01

Start with arithmetic, not with tools

Divide the file size by the page count. That single number immediately tells you which kind of problem you have. A text document exported from a word processor typically lands in the low tens of kilobytes per page. A page carrying a reasonable photograph might be a few hundred. If your figure is in the megabytes per page, the cause is image data and you can go straight to that suspect. If the figure is enormous but the page count is tiny — a three-page file weighing dozens of megabytes — something non-page-shaped is likely riding along inside.

Then look at whether the pages are scans or generated content. Zoom in hard on a letterform. If the edges stay clean and geometric, the text is real text and the page is generated. If it breaks into grey pixels and speckle, you are looking at an image of a page, and the whole diagnosis changes: everything you see is one big raster stream, and its resolution and colour mode are the only things that matter.

Finally, check whether the size is spread evenly or concentrated. If a document is mostly modest pages with two or three enormous ones, you have a local problem — a couple of photographs or a pasted screenshot at native camera resolution — and a targeted fix will keep the rest of the document untouched. Splitting the file and comparing the sizes of the parts is a crude but extremely fast way to find the heavy pages when nothing else is to hand.

02

The suspects, and the evidence each one leaves

High-resolution embedded images are the usual culprit, and they leave an obvious fingerprint: the heavy pages are the ones with pictures on them. The telling detail is not the image’s file size but the mismatch between its pixel dimensions and the space it occupies on the page. A photograph from a modern phone is several thousand pixels wide; dropped into a report as a two-inch illustration, almost all of those pixels are unused, and they are being stored and transmitted anyway.

Fonts leave a different signature: the file is large but the size is spread flat across pages, and the burden does not move when you touch the images. Embedding a font that was not subset stores the entire typeface rather than the glyphs the document uses, and for a CJK font with tens of thousands of glyphs that can dominate a short document completely. Several weights of the same family, each embedded in full, multiply the problem. Duplicated resources look similar: the same logo or background image stored once per page instead of shared as a single object, so the file grows in proportion to page count even though nothing visibly changes from page to page.

Then there are the invisible passengers. Retained incremental-update history is the sneakiest: PDF permits changes to be appended rather than rewritten, so a file that was edited and saved twenty times can still contain every earlier state, including content you believed you had deleted. The symptom is a file far larger than its visible content can explain, often with a long editing history behind it. Embedded file attachments behave the same way — a spreadsheet attached to an invoice is simply carried inside the PDF at its full size. Uncompressed content streams, produced by generators that never applied a filter, show up as a file whose size drops sharply under a purely lossless optimisation pass. And a scanned page of plain black text stored as a full-colour image is paying for three colour channels to describe something that is ink or paper and nothing else.

03

Turning suspicion into evidence

An inspector that reports the file’s internal composition settles most cases in one step, by listing the objects and showing what share of the total each category of stream occupies. When images dominate the breakdown, list them and compare each one’s pixel dimensions against the size it is drawn at. When fonts dominate, list the embedded fonts and check whether each is subset — a name without the six-letter subset prefix, attached to a large font program, is your answer. A colour inspector tells you whether pages nominally black and white are actually being stored in colour.

A few checks need no tooling at all. To test for retained revision history, open the document and use Save As, or export a fresh copy, rather than saving in place; a rewrite that discards prior revisions can collapse the size dramatically, and if it does, history was your problem. To test for duplicated per-page resources, note whether size grows roughly linearly with page count in a document whose pages share a common background or letterhead. To test for attachments, look in the viewer’s attachments panel, which most people never open.

Only once you know the answer should you choose a remedy, and then choose the narrowest one that addresses it. History and unreferenced objects go away with a lossless rebuild, costing nothing visible. Duplicated resources are fixed by sharing rather than by degrading anything. Oversized images want resampling to match their placed size, which is lossy but targeted. A colour scan of monochrome text wants converting to greyscale or bilevel. Reaching for a maximum-strength compression preset before diagnosing is how documents end up both smaller than necessary and worse-looking than necessary.

Worth repeating

MyPDFilles tools described here run on your device.

When an article refers to a MyPDFilles tool, its parsing, compression or recognition runs in JavaScript and WebAssembly inside your browser tab on bytes read from your disk. Educational references to external software are not covered by that claim; review the external provider’s own privacy and security information.

How to verify MyPDFilles processing

Questions on this topic

My PDF is huge but every page is just text. How?

Three explanations cover nearly all such cases: the pages are actually images of text rather than text, a large font was embedded without subsetting, or the file is carrying retained revision history from repeated saves. Checking whether the text is selectable distinguishes the first from the other two immediately.

Why did the file shrink enormously when I just re-saved it?

Because a full rewrite discards what an incremental save keeps. Appending changes rather than rebuilding the file means earlier revisions stay inside it, and a document edited many times can be mostly history. Exporting a fresh copy writes only the current state, which is why the saving can be so large.

Deleting pages barely reduced the size. Why not?

Removing a page removes its reference, but the objects behind it may still sit in the file unless the writer rebuilt it and dropped what nothing points to any more. An optimisation pass that discards unreferenced objects is what actually reclaims the space.

Is a big PDF ever simply unavoidable?

Sometimes, yes. A long, high-resolution colour scan intended for archival or print genuinely needs its pixels, and compressing it hard would defeat its purpose. The right move is then to pick a target for the destination — one copy at archival resolution, a separate lighter copy for email.