A machine reading of your page, with mistakes in it. Not a verified transcription. Optical character recognition guesses each character from its shape. A wrong digit looks exactly like a right one, and the engine can be confident about a word it has misread. Proofread against the original before you rely on any of it. Used at your own risk. Open for the full scope limits.
What this tool actually does
It turns each page into a picture and asks the Tesseract engine what the marks on it say, in English. Nothing on the page knows what the document is, what a sensible value would be, or whether a sentence makes sense.
- A PDF page that already holds text is not read from pixels. A page with 20 or more characters of its own text is copied from the file, and its row says so. That text is whatever the file's author or an earlier OCR pass put there. A scan with a typed header or a stamp over an image body counts as a text page and its body is skipped: tick "Read every page from pixels" to read the body.
- Pages are drawn at 300 DPI or less. A very large page is drawn at a lower resolution to fit in memory, and an image over 4,000 pixels on its longest side is scaled down. The row gives the resolution used.
- Four repairs can be made before reading, and each is stated in the page's row. Paper that measures brighter in some places than others, as under a shadow, is leveled to plain gray ("lighting evened"). Print tilted between 6 and 30 degrees is turned by that angle ("turned"). Print under about 20 pixels tall, as on a screenshot, is enlarged two or three times ("enlarged"). A page that reads as nonsense is read again upside down, and that reading is kept when it is clearly the better one ("turned the right way up"). A page that measures level, straight and large enough is read exactly as it came. All four were measured on test pages drawn by a program, not on photographs, and any of them can misjudge a page that is mostly pictures or ruled lines.
- A second engine reads a page the first one read badly. When Tesseract scores a page under 70, or finds nothing on a page that holds print, the page is read again by PP-OCR, a different kind of reader. Its reading replaces the first only when it scores 90 or more on its own scale and found at least half as much text, and the page's row says which engine the text came from. A page whose lines slope more than 30 degrees is turned level for it first, whichever way up it lies. The second engine is about 27 MB, fetched from this site the first time it is needed, and it is slow: a full letter page took it 16 seconds on a fast computer and 43 seconds with one worker, which is what a phone gets. Set "Second engine" to "Never" to keep it out. It was tried only on test pages drawn by a program.
- Two engines agreeing is not proof. "Compare both on every page" shows the two readings side by side with the words they read differently marked. A marked word is worth a look. An unmarked one was read the same way twice, and both readers confuse the same look-alike characters.
- Confidence is the engine's opinion of itself. It is a score from the same model that did the reading, not a measurement against the truth. The two engines score on different scales, and their figures are never compared with each other.
What it gets wrong
- Characters that look alike. 0 and O, 1 and l and I, 5 and S, 8 and B, a comma and a period, and the letters "rn" read as "m". In a test of this page's own engine, "rn m" came back as "mm" with a confidence of 87 out of 100.
- Numbers above all. A word that is misread usually stops being a word, which a reader notices. A misread digit is still a number. Amounts, dates, doses, part numbers, account numbers and serial numbers must be checked one by one.
- Reading order. The engine decides what is a column and reads each one top to bottom. Columns, forms, sidebars, footnotes, headers and captions come out in the order it guessed, which may not be the order a person reads them in. A table with ruled lines is usually read row by row, and a table without them is usually read down each column, so a figure ends up far from its label. Table cells lose their rows and columns either way.
- Dropped and invented text. Faint print, small print, text over a picture, and text near a fold, a shadow or the edge of a photo can be left out with no sign that anything was there. Specks, rules and logos can be read as characters.
- A word that is not flagged is not a word that is right. "Check these" lists the words the engine doubted, each beside the pixels it was read from. It says nothing about the rest.
The searchable PDF
- The hidden text is the same reading, mistakes included. Searching the saved file for a word or a number that was misread finds nothing, and a search that finds nothing is not proof the document does not contain it.
- It is not redaction and not a cleaned copy. The page image is carried over unchanged, with everything that was visible on it.
- Only pages read from pixels are given text. Pages that already held text are left exactly as they were, including any wrong text an earlier program put there.
- Rewriting a PDF can drop things. A digital signature stops validating once the file is changed, and forms, bookmarks or attachments in an unusual file may not survive. Keep the original.
- An image is given a page size by assumption. A picture carries no physical size, so its page is sized as if scanned at 300 DPI.
Outside what it reads
- Handwriting, signatures and anything written by hand on a printed form.
- Languages other than English, and accented or non-Latin characters, which come out as the nearest English shape.
- Mathematics, chemical formulas, music, barcodes and QR codes.
- Photos taken from the side so the page narrows away from the camera, out of focus, or of a curved page.
- A page tilted more than 30 degrees with "Second engine" set to "Never". Only the second engine reads one.
- Print under about 9 pixels tall, which is still misread after it is enlarged.
- Password-protected PDFs, which are refused.
Never use this as the only transcription of
- A prescription, a medicine label, a dose, a lab result or any clinical record.
- A contract, a court filing, an identity document or any legal record.
- An invoice, a bank or tax document, or any amount of money.
- A drawing, a specification, a nameplate or a safety data sheet whose numbers someone will build or act on.
Your file stays in this browser
Both reading engines, their language data and the PDF library are served from this site and run on your device. The file you choose is not sent to this site or to anyone else, and nothing from it is stored after you leave the page. The page fonts are loaded from Google Fonts, which sees a request from your browser but nothing about your file. That is a description of how the page is built, not a security guarantee: do not put a document here that policy, contract or law says must not be opened in a web browser.
Before you rely on the text
Read it against the original, start with the pages and words marked for checking, and then check every number whether it was marked or not. Where the text is going into a record, have a second person compare it.
Source
A PDF, an image, or a photo. It is read on this device and never uploaded.
Drop a file here, paste an image, or
PDF, PNG, JPEG or WebP. Up to 25 pages and 50 MB. English text. "Take a photo" opens the camera on a phone and a file picker on a computer. For a photo, hold the camera square over the page.
Off, a PDF page that already holds text is copied from the file, which is exact when that text is good. Turn it on when the copied text is wrong or missing part of the page.
The first engine is Tesseract. A page it scores under 70, or finds nothing on, is read again by a second engine, PP-OCR, which is fetched from this site when first needed (about 27 MB). Comparing reads every page with both and shows where they differ.
Text
Nothing read yet.
The searchable PDF is your original page with the reading laid over it as hidden text, so it can be searched and copied from. The page itself is not changed, and the hidden text carries every mistake in the reading.
This is a machine reading, not a checked transcription. A misread digit looks exactly like a correct one, the engine can be confident and wrong, and a word that is not marked for checking is not a word that is right. Proofread against the original, numbers first. Used at your own risk. See the full disclaimer at the top of the page.
Check these
| Page | Pixels | Read as | Confidence |
|---|