14 articles · 95 min in total
Understand the format, and the tools stop surprising you.
These are explanations rather than instructions. What a PDF is at the level of its bytes, why lossy compression ruins text, what document encryption genuinely protects — the mechanisms underneath the buttons.
If you want the click path for a specific job, the guides are the right place. These pages answer the question behind the job — which is usually what makes the difference between a result that works and one that merely completed.
01 · 2 articles
Format
- 8 min read3 parts
What is a PDF?
A PDF records marks on a fixed canvas rather than the structure of a document. Understanding that one decision explains almost every PDF frustration.
- 7 min read3 parts
What is PDF/A?
An archival subset of PDF that forbids anything whose appearance depends on the outside world. Same format, deliberately fewer freedoms.
02 · 2 articles
Text
- 7 min read3 parts
What is OCR?
Optical character recognition is pattern matching on images of glyphs. It produces a guess, and the quality of that guess is decided long before the software runs.
- 6 min read3 parts
What is a searchable PDF?
A two-layer artefact: the original page image you see, plus invisible text underneath it that search and selection actually operate on.
03 · 4 articles
Security
- 7 min read3 parts
What is PDF metadata?
Every PDF carries fields describing itself. Most were filled in by software you never configured, and some of them say more than you intended.
- 7 min read3 parts
What is PDF encryption?
Real cryptography applied to the contents of a PDF. Without the key the bytes are unreadable — which is what separates it from permission flags.
- 6 min read3 parts
What are PDF permissions?
A set of bits recorded in the file asking software to withhold certain actions. Honoured by convention, not enforced by cryptography.
- 7 min read4 parts
Digital vs electronic signature
One is an image placed on a page. The other is a cryptographic proof that the bytes have not changed since signing. The words sound interchangeable and the guarantees are not.
04 · 3 articles
Quality
- 8 min read4 parts
How PDF compression actually works
A PDF is not compressed as a whole. Each stream inside it carries its own filter, and knowing which filter is doing what explains every size-versus-quality trade-off.
- 7 min read3 parts
Why is my PDF so large?
A diagnostic walkthrough rather than a list of tips: how to work out which part of a specific file is consuming the bytes before you change anything.
- 6 min read3 parts
What is DPI?
DPI is not something an image has — it is pixels divided by physical inches. Once you see it as a ratio, resolution decisions become arithmetic instead of guesswork.
05 · 3 articles
- 6 min read3 parts
A4 vs Letter paper size
Two standards, a few millimetres apart in each direction, and no clean way to convert between them. The differences are small and the consequences are not.
- 6 min read3 parts
Raster vs vector in a PDF
Every PDF is a hybrid document. Knowing which parts of a page are pixels and which are equations predicts exactly how it will behave under zoom and in print.
- 7 min read3 parts
PDF fonts explained
Embedding puts the typeface inside the file; subsetting keeps only the glyphs used. When neither happens, the viewer substitutes — and your layout moves.
Why bother with the theory
Most PDF frustration is a mismatch, not a bug.
A PDF records where marks sit on a fixed canvas. It does not record paragraphs that flow, which is why editing a sentence does not push the rest of the text along, why a scanned page contains no words until OCR invents some, and why “convert to Word” can only ever be a reconstruction rather than a translation.
Once that single fact is in place, most surprising tool behaviour stops being surprising. These articles exist to put it in place — with enough specificity that you can predict the outcome before you run anything.