PDF to Excel
Convert PDF tables to editable Excel spreadsheets (.xlsx).
Drop a PDF with tables
or click to browse
Convert PDF tables to editable Excel spreadsheets (.xlsx).
or click to browse
.xlsx. Each page becomes its own sheet, so nothing is silently merged.The document is never uploaded. Detection, recognition and the spreadsheet build all run inside this browser tab.
Pulling text out of a PDF is easy. Working out which piece of text belongs in which cell is the actual problem, and it is where most converters fall over — you get a spreadsheet with everything crammed into column A, which is no more useful than the PDF was.
So the page is analysed for structure rather than skimmed for words. Ruling lines are detected where a table has them, and column and row boundaries are inferred from the alignment of the text itself where it does not — which matters, because plenty of real invoices and statements are laid out with whitespace and no borders at all. Merged header cells, where one heading spans several columns, are recognised rather than smeared across the row.
When a page holds no table at all, its text is laid out in reading order instead of being discarded, so a document that mixes narrative pages with tabular ones converts completely rather than losing half of itself.
A scanned page has no text to extract, so it gets a different treatment: the table grid is located visually, and then each cell is read individually.
That per-cell approach is the reason the output lands in the right columns. The obvious alternative — recognise all the words on the page, then try to group them into cells afterwards — falls apart on exactly the documents people care about, because a slight skew or a tight column gap sends a figure into the neighbouring column, and a misplaced number in a financial table is worse than no number at all. Reading a cell at a time means each value is anchored to the cell it came from by construction.
It is more thorough than a plain page-level scan, so allow roughly 40 to 55 seconds per scanned page on a typical laptop. Pages that already carry a text layer convert in a moment; the time is spent only where it buys accuracy. For a long scanned document, start it and leave the tab open.
Two things noticeably improve results: straighten sideways pages with Rotate PDF first, and pick the right language, including a bilingual combination for the Indian government and legal documents that routinely mix English with a regional script.
Reliably good: bank and credit card statements, invoices and purchase orders, price lists, inventory and stock reports, financial statements, timetables, attendance registers, survey results, and product catalogues. Anything with a genuine row-and-column structure, whether or not it is ruled.
Needs a look before you trust it: tables split across a page break, where the continuation may start a new sheet; deeply nested headers three levels down; cells containing several lines of free text; and pages where two tables sit side by side. These convert, and a quick scan of the output is worth the thirty seconds.
The general rule is that the more a page looks like a spreadsheet, the better it converts. A financial statement comes across almost perfectly. A magazine page with a chart embedded in an article is doing something else, and PDF to Word is usually the better route for that — it rebuilds ruled tables as real Word tables inside a flowing document.
Yes, and it is built for them. The table grid is located visually and then each cell is read on its own, which is what keeps figures in the right columns — the failure mode of page-level recognition. Set the language before converting, and allow roughly 40–55 seconds per scanned page.
Values come across as data rather than as styled cells, so apply currency, date or percentage formatting in Excel afterwards. That is generally what you want, since formatting depends on your locale rather than the document's, and it means the underlying figures are clean and ready to calculate with.
Each page becomes its own sheet, which keeps page boundaries visible and stops a repeated header row being mistaken for data. When you want one continuous table, copying the sheets together in Excel takes a moment — and doing it afterwards means you can see exactly where each page ended.
Digits are read as they appear, so a value written in the Indian grouping style comes through with its digits intact. Apply the Indian number format in Excel afterwards to display it in the convention you want. Bilingual language options — English + Hindi, Tamil, Telugu, Marathi — are available for scanned documents that mix English with a regional script.
Its text is laid out in reading order on that page's sheet rather than being dropped, so a document mixing narrative and tabular pages converts in full. You can delete the sheets you do not need.
Use Split PDF in Extract specific pages mode to pull out the pages holding the tables, then convert that. For a long scanned document this is worth doing — it is the quickest way to cut the conversion time.
Remove the password first with Protect PDF, then convert. Encryption exists precisely to stop the page content being read.
No. Table detection, cell recognition and the spreadsheet build all happen inside this browser tab, and the .xlsx goes straight to your downloads. Bank statements and financial records never leave your machine — go offline before you start and it still works.
Privacy: The PDF is parsed with PDF.js and the Excel file built with SheetJS — both run inside this browser tab. Financial documents never leave your device.