Skip to content

Extract From PDF

PDF Text Statistics

One report that describes how a document is written, not just how long it is.

Processed locally in your browser. Your file is not uploaded to any server. How this works

  1. 01Add your file
  2. 02PDF Text Statistics
  3. 03Download

Runs in your browser · nothing uploaded

About PDF text statistics

Where the single-metric tools answer "how much", this one answers "what kind of writing is this". After extraction the text is segmented into sentences using abbreviation-aware punctuation rules, so "Dr. Chen" and "e.g." do not falsely end a sentence. From those sentences it derives average sentence length, the share of sentences over thirty words, average syllables per word estimated by vowel-group counting, and from those a Flesch reading-ease score and its approximate grade level. It adds a type-token ratio for vocabulary variety, the twenty most frequent content words after stop-word removal, and characters-per-page as a density measure that tells you whether a forty-page document is dense or airy. Treat readability scores as rough instruments: they measure sentence and word length, not clarity, and technical prose with short sentences can score easy while being impenetrable.

How to PDF text statistics

  1. 01

    Load the document

    Drop in the PDF you want profiled. All analysis is local.

  2. 02

    Start with the summary

    Sentence length, readability grade and density give you the shape of the document in three numbers.

  3. 03

    Look at frequency

    The top content words reveal what the document is actually about — often the fastest way to triage an unfamiliar report.

  4. 04

    Download the report

    Keep it alongside the document, or use it to justify an editing pass.

What this tool does

  • Abbreviation-aware sentence segmentation, not a naive split on full stops
  • Flesch reading-ease score with approximate grade level
  • Type-token ratio as a vocabulary-variety measure
  • Top twenty content words with stop words removed
  • Characters-per-page density and a long-sentence share to target editing

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • Readability formulas were calibrated on English prose; scores for other languages are indicative at best.
  • Syllable counts are estimated from vowel groups, so unusual names and loanwords can skew the grade slightly.
  • Scanned PDFs must be OCR'd first — with no text layer there is nothing to analyse.
  • Tables, references and code blocks distort sentence-length statistics; exclude those pages for a cleaner read.

How your file is handled

This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about PDF text statistics

Which readability score do you report?

Flesch reading ease, plus the grade level it implies. It is the most widely recognised measure, which makes it the most useful to quote even though every such formula is an approximation.

What does the type-token ratio tell me?

How varied the vocabulary is: unique words divided by total words. It falls naturally as documents get longer, so compare it between documents of similar length rather than treating it as absolute.

Can I use this to check a document against a plain-language requirement?

As a first pass, yes — it will flag long sentences and a high grade level. Formulas cannot judge jargon or structure, so a human review is still required.

Is any of this sent anywhere for analysis?

No. Segmentation, scoring and frequency counting all run in JavaScript in your tab on text that never leaves your machine.

Read more about this