PDF to Markdown
Headings, lists, tables and images turned into clean Markdown, with an optional bundle of chunks for a knowledge base. All of it in your browser.
Open in PDF ARENADrop your PDF here or click to open
or click to select files
Reading order first
Two columns, sidebars and running headers get sorted out before a single character is written.
Tables with a verdict
Each table carries a confidence figure. A shaky one is flagged, never quietly invented.
A package, not a dump
Chunks with their heading path, pages and hash. Drop it straight into a knowledge base.
How to convert a PDF to Markdown
Open your PDF
Drop it in. Conversion starts straight away and the file stays on your machine.
Read what came out
The result lands in a box you can edit. Fix a heading, join a split sentence, delete a stray line.
Pick your flavour
GitHub, CommonMark, MDX, Obsidian or Notion. Images in a folder or inline. Grids as Markdown, HTML or CSV.
Download, or build the package
One .md, or a ZIP holding chunks, pictures, grids and provenance for a knowledge base.
What gets reconstructed, and what gets flagged
Columns are found by hunting gutters: vertical strips no line crosses, spanning most of the text band. Width alone proves nothing. A single-column report with one two-cell grid has narrow gaps scattered everywhere, and a threshold on width would call each of them a column boundary, so what settles it is that a real gutter is one every body line respects and the share of height covered on both flanks decides. Running banners are matched with digits normalised away, since "Page 3 of 40" and "Page 4 of 40" are one footer. Six pages out of ten. Below that, the line stays. Grids are the honest part. Three or more consecutive lines whose fragments begin at roughly matching x positions form a table; prose does not. Boundaries cluster with a tolerance proportional to type size, and every table emerges carrying a confidence figure: the share of its lines that genuinely respect them. Below 0.6 it still gets written, flagged and raised as a warning. Inventing cells would be worse. Pictures come cropped from the rendered page at 216 dpi rather than pulled out of the original stream. That costs native resolution, and buys survival across every encoding a PDF can carry. The bundle holds one chunk per line in JSONL, each with its heading path, pages, token estimate and SHA-256, plus a map naming which page produced each block. Chunks break on block edges. Always. Splitting a paragraph mid-sentence to hit a token count yields fragments nobody reads alone, which defeats retrieval entirely.
Why this one
Reading order solved first: columns, sidebars, repeating banners.
Tables carry a confidence figure, and a doubtful one says so.
Chunks, provenance and hashes in one bundle, assembled on your machine.
Questions about PDF to Markdown
Short answers, limits included
Does it handle two-column documents?
Are the tables reliable?
What happens to the images?
What is in the RAG package?
Does it convert formulas to LaTeX?
Updated on
Do more with your PDFs
Discover all the professional tools we have prepared for you. All in one place.