Skip to content

Quality · 8 min read · Updated 2026-09-19

How PDF compression actually works

A PDF is not compressed as a whole. Each stream inside it carries its own filter, and knowing which filter is doing what explains every size-versus-quality trade-off.

People talk about compressing a PDF as though the file were a ZIP archive — one algorithm applied to one blob. It is nothing like that. A PDF is a collection of numbered objects, and compression is declared individually on each stream object inside it, with a filter chosen to suit the kind of data that stream holds. Understanding that structure is what turns compression from a slider you drag into a decision you can reason about.

How PDF compression actually works

01

Compression is declared per stream, not per file

The unit of compression in a PDF is the stream: a run of raw bytes attached to a dictionary that describes it. A page’s drawing instructions are a stream. Each embedded image is a stream. Each embedded font program is a stream. So is an ICC colour profile, an attached file, and the XMP metadata packet. Every one of those dictionaries may carry a /Filter entry naming the decoding filter a reader must apply before the bytes make sense.

Because the declaration is per stream, a single PDF routinely mixes several compression schemes at once. The text of page four might be deflated, the photograph on it stored as JPEG data, the scanned signature beside it stored as a bilevel image under an entirely different filter, and the embedded font compressed with a fourth. There is no global setting that governs all of them, which is why two files that look identical on screen can differ wildly in size: they made different per-stream choices.

Filters can also be chained. A dictionary may list an array of filters, applied in order, most commonly when a stream is deflated and then encoded into a text-safe representation for transport. The important consequence for anyone trying to shrink a document is that the question is never "how compressed is this PDF" but "which streams are large, and what filter is each of them using".

02

The filters, and what each one is actually for

FlateDecode is the workhorse. It is the zlib/DEFLATE algorithm, exactly as used in ZIP and PNG, and it is completely lossless — decoding returns the original bytes byte for byte. It is what page content streams, font programs, metadata and most structural data use, because those are symbolic data where losing a byte would mean losing meaning. Text in a PDF is essentially always compressed with Flate, which is why text-only documents are small and why nothing you do to image settings makes a text-only file meaningfully smaller.

DCTDecode is JPEG. The name comes from the discrete cosine transform at the heart of the algorithm, and a DCTDecode stream contains a JPEG bitstream that the reader hands to a JPEG decoder. It is lossy by design, discarding the high-frequency detail human vision is least sensitive to, and it is the right filter for photographs and other continuous-tone imagery. JBIG2Decode and CCITTFaxDecode serve the opposite kind of image: bilevel content, one bit per pixel, black or white and nothing between. Both are built around the statistics of scanned text and line art, and on a clean bilevel scan they are dramatically more efficient than storing the same page as a photograph.

Two more appear mostly in older or unusual files. LZWDecode is the legacy predecessor of Flate, lossless and still supported for compatibility but no longer chosen by modern producers. RunLengthDecode simply collapses runs of identical bytes; it is trivially simple, occasionally useful for very flat synthetic imagery, and otherwise rare. Alongside the stream filters, two structural features compact the file itself rather than its content: object streams, which pack many small objects into one compressed stream instead of leaving each exposed, and cross-reference streams, which replace the plain-text cross-reference table with a compressed one. Neither touches a single pixel — they reduce the overhead of the file’s own bookkeeping.

03

Which operations lose information, and which do not

Sorting compression operations into lossless and lossy is the single most useful distinction to hold onto, because it tells you what can be undone. Stripping metadata is lossless with respect to the page: the visible document is untouched. Re-deflating streams that were stored uncompressed or under a weaker filter is lossless. Rebuilding the file’s structure — writing object streams and a cross-reference stream, discarding orphaned objects that nothing references any more, sharing a resource that had been duplicated — is lossless. All of these reduce bytes without altering what is drawn on the page.

Two common operations are unambiguously lossy. Downsampling an image throws away pixels: a 3000-pixel-wide scan resampled to 1200 pixels has lost two thirds of its samples and no later step can reconstruct them. Re-encoding an image as JPEG, or re-encoding an existing JPEG at a lower quality setting, quantises detail away permanently, and because the input was already an approximation the second pass compounds the first. Converting colour to greyscale is lossy in the same irreversible sense: the hue information is simply gone.

This is why a sensible compression pass is ordered. Do the lossless work first and measure what it achieved, because on some files — particularly ones exported by careless generators, or carrying a long revision history — the lossless steps alone are enough. Only then reach for downsampling and re-encoding, on the specific images that justify it, and keep the original file. A lossy pass is a one-way door, and the version you compressed is the version you will have.

04

DPI and JPEG quality are independent axes

The most persistent misunderstanding about image compression is that resolution and quality are the same dial. They are not, and conflating them produces both bloated files and ugly ones. Resolution is how many pixels exist. JPEG quality is how coarsely each of those pixels’ frequency coefficients is quantised when encoded. You can vary either while holding the other fixed, and the two failure modes look completely different.

A large image at low quality is soft and blocky at full size, with visible ringing around sharp edges, yet still occupies more bytes than it should for how bad it looks. A small image at high quality is crisp within its own dimensions but goes visibly fuzzy the moment it is enlarged or printed at size. Neither is a good bargain. The useful approach is to set resolution from the destination — what the image will physically measure on the page, and whether the page is going to a screen or a press — and then choose a quality setting high enough that artefacts are not visible at that size.

Resolution decisions also interact with the choice of filter. A scanned page of black text stored as a full-colour JPEG is paying twice: once for colour channels carrying no information, and again for a lossy filter that handles sharp letterforms badly. The same page stored as a bilevel image under a filter designed for bilevel content is usually smaller and sharper at the same time. Choosing the right filter for the content type frequently beats turning any quality slider down.

Worth repeating

MyPDFilles tools described here run on your device.

When an article refers to a MyPDFilles tool, its parsing, compression or recognition runs in JavaScript and WebAssembly inside your browser tab on bytes read from your disk. Educational references to external software are not covered by that claim; review the external provider’s own privacy and security information.

How to verify MyPDFilles processing

Questions on this topic

Is compressing a PDF the same as putting it in a ZIP?

No. Zipping wraps the whole file in one more layer, and since a PDF’s large streams are already compressed there is usually almost nothing left to squeeze — the ZIP comes out a few per cent smaller at best. Compressing a PDF means changing the filters and data of the streams inside it, which is an entirely different operation.

Can I compress a PDF without losing any quality at all?

Often, yes, up to a point. Stripping metadata, re-deflating uncompressed streams, discarding unreferenced objects, sharing duplicated resources and rebuilding the cross-reference structure are all lossless — the page renders identically afterwards. How much they save depends entirely on how wasteful the original producer was.

Why did compressing my document barely change its size?

Almost certainly because it is mostly text. Text and vector artwork are stored in Flate-compressed content streams that are already close to their efficient size, so there is little headroom. Meaningful reductions come from image data, and a document with no large images has nothing much to give up.

Does compressing twice make a file smaller still?

Not usefully, and it can do harm. A second lossless pass finds almost nothing the first one missed. A second lossy pass re-encodes images that are already approximations, degrading them further for a small saving. If a file is still too large after one careful pass, change the settings rather than repeating the pass.