PDF to Text
Extract all text content from your PDF file.
Drop a PDF to extract its text
or click to browse
Extract all text content from your PDF file.
or click to browse
The document is read inside this browser tab. Nothing is uploaded, which is worth knowing given that the PDFs people extract text from tend to be contracts, statements and reports.
A PDF does not store paragraphs. It stores instructions of the form "draw these characters at this coordinate", in whatever order the producing software found convenient — which is frequently not reading order. Ask for the raw items and you can get a page's footer before its heading, or a two-column layout zippered together line by line.
So the text is reassembled geometrically. Every text fragment is taken with its position on the page, sorted from the top down, and grouped into lines: fragments whose vertical positions sit within about five points of each other belong to the same line, which is tight enough to keep adjacent lines apart and forgiving enough to tolerate the slight baseline variation you get with mixed font sizes and subscripts. Within each line, fragments are ordered left to right and joined, and runs of whitespace are collapsed.
The result reads the way the page reads. Headings arrive before the paragraphs under them, lines stay whole, and words that a PDF had split into three separate drawing operations come back as one word.
Each page is introduced by a --- Page 3 --- marker, which sounds like a small thing until you are searching a 200-page extraction for the clause you need and want to know which page to open. Pages with no text on them are skipped rather than producing an empty heading.
A scan holds pictures of words, not words, so there is no text layer to pull out. Rather than handing you a file with nothing in it, the tool checks and tells you, then offers the route that works: run the document through OCR first. That reads the page images and writes a genuine text layer into the PDF — and once it has, extraction here returns the full text exactly as it would from a born-digital document.
OCR is also the answer for photographed pages, faxes, and any document that arrived as an image inside a PDF wrapper. It handles a wide range of languages and scripts, and the searchable PDF it produces is useful well beyond this one tool: PDF to Word, Edit PDF and in-document search all start working on it.
Plain text is the right answer when the words are what you want and the formatting is not. When the destination needs more, other tools take it further from the same source document.
Editable document with headings, lists and tables — PDF to Word rebuilds paragraph structure, maps headings onto Word styles, keeps bold and italic inside sentences, and reconstructs ruled tables as real tables.
Numbers destined for a spreadsheet — PDF to Excel targets the cell grid directly and gets columns into columns, which plain text extraction cannot do by design.
Changing a few words and keeping the PDF — Edit PDF lets you click a word in the document and retype it, matching the original font, size, weight and colour. No conversion round trip.
Slides from a PDF deck — PDF to PowerPoint produces editable slides rather than a wall of text.
Almost always that the PDF is a scan — pictures of pages, with no text layer underneath. The tool detects this and points you at OCR, which reads the page images and writes a real text layer into the document. Bring that searchable PDF back here and the full text comes out.
Table content comes through row by row, in reading order, since plain text has no notion of a cell. When the table structure is the point, PDF to Excel extracts it into a spreadsheet grid, and PDF to Word rebuilds ruled tables as genuine Word tables.
Every page is extracted in one pass, each marked with a --- Page N --- header, so finding the page you want in the output is quick. To work with a single page as a file, Split PDF pulls it out first and you can extract from that.
UTF-8, so accented characters, currency symbols, and non-Latin scripts survive intact. Any modern editor opens it correctly. If a much older Windows tool shows odd characters, re-saving the file as UTF-8 with a BOM in a text editor usually settles it.
Extraction follows the geometry of the page, so unusual layouts — text inside graphics, rotated headings, heavily overlapping frames — can order differently to how the eye reads them. Text drawn as vector outlines rather than characters is artwork rather than text and is not part of the text layer at all; running the page through OCR reads it visually and recovers it.
Remove the password first with Protect PDF, then extract. Encryption exists precisely to stop the content stream being read.
No. Text extraction is light work compared with rendering pages, so documents in the hundreds of pages process quickly, including on a phone. The progress bar reports the page it is on throughout.
No. The document is parsed by PDF.js inside this browser tab, and the text never leaves it — the copy goes to your clipboard and the download goes to your downloads folder. Disconnect from the network before you start and it still works.
Privacy: The PDF is parsed entirely in your browser with PDF.js. Nothing is sent over the network — the extracted text exists only in this tab until you download or copy it.