Scanned & image-only PDFs

Convert scanned PDF to Markdown with OCR

Drop an image-only or scanned PDF and get clean, selectable Markdown. Built-in OCR works across many languages, rebuilds real tables and keeps formulas – no account, no separate OCR step.

Short answer

Yes – a scan becomes selectable Markdown

A scanned PDF is just images of pages, so plain copy-paste returns nothing or garbled characters. PDF to Markdown detects image-only pages and runs OCR (optical character recognition) automatically, turning the pictures of text into real, selectable Markdown – headings, lists, tables and all. It works on documents scanned in many languages, including mixed-language pages, and you can convert in the browser without signing up.

How to

Convert a scanned PDF in 4 steps

No account needed. OCR runs automatically, or force it when a PDF has a bad text layer.

1

Open the converter

Install the Chrome extension or open the web app. Both work anonymously.

2

Add the scanned PDF

Drag in the file, pick it from disk, or paste a direct PDF URL. OCR runs automatically on image-only pages; toggle force OCR when the existing text layer is wrong.

3

Wait for the job

Status goes queued, processing, ready. OCR is heavier than reading digital text, so scans take longer than native PDFs.

4

Copy or download

Preview the rendered Markdown and the raw source, then copy it to your clipboard or download a .md file.

Tip: automating bulk scans? Skip the UI and call the REST API or hosted MCP – same OCR, driven from your own code or agent.

What OCR preserves

More than just plain text

Recognizing characters is the easy part. The converter rebuilds the document structure a scan loses, so the Markdown is usable by people and models alike.

Many languages

Reads scans across many languages, including mixed-language pages, into selectable text.

Real tables

Scanned columns become genuine Markdown tables instead of a jumble of misaligned lines.

Formulas kept

Mathematical notation is preserved rather than flattened into garbled characters.

Force OCR

Override a bad or partial text layer and re-read the page images when the embedded text is wrong.

Links & footnotes

Where present, hyperlinks and footnotes carry over as Markdown links instead of being dropped.

Engine choice

Convert with MinerU or Docling, depending on the document and the result you want.

Quality in, quality out

What affects OCR accuracy

OCR reads pictures of text, so the cleaner the scan, the cleaner the Markdown. A few things make the biggest difference.

Best results

Sharp, high-resolution pages. Around 300 DPI or higher, with text that is crisp rather than blurry.
Straight, deskewed scans. Pages that are upright, not rotated or warped.
Good contrast. Dark text on a light, even background, without heavy bleed-through.

Harder for OCR

Faint or low-resolution scans. Copies of copies, or small screenshots, lose the detail OCR needs.
Skew, shadows and clutter. Phone photos at an angle, page curl or busy backgrounds reduce accuracy.
Handwriting. The engines target printed text; handwritten notes are not reliably recognized.

What you can convert: any PDF up to the size limit, including image-only, mixed digital and scanned documents, multi-column layouts and tables. Output is a single Markdown file, or the raw Markdown text over the API.

Troubleshooting

Common problems and quick fixes

Garbled or wrong text

The PDF has a bad embedded text layer. Turn on force OCR so the converter re-reads the page images instead of trusting that layer.

Nothing recognized

Usually a very faint or rotated scan. Rescan straight at a higher resolution, or improve contrast, then convert again.

Result marked truncated

A long scan hit the time budget and was returned partially. Split the document into smaller files, or use a paid tier with a longer budget.

Messy tables

Try the other engine: MinerU is robust on scans and complex layouts, while Docling is fast on clean, simple pages.

A recognized table comes back as real Markdown, ready to paste or index:

| Quarter | Revenue | Growth |
| ------- | ------- | ------ |
| Q1      | $1.2M   | +8%    |
| Q2      | $1.4M   | +17%   |
What to expect

Free tier limits & long scans

Free tier limits

Active slots (queue depth)3
Max PDF size10 MB
Time budget per document15 min
Ready result retention1 hour

Paid tiers raise every limit and add a longer time budget for heavy scans. Compare plans →

Long or low-quality scans

Partial results are flagged. If a long scan hits the time budget, you get what was processed, marked truncated, instead of an error. Split the file or use a longer paid budget.
Legibility matters. OCR accuracy follows the scan: a clean, straight, reasonably high-resolution page reads far better than a faint or skewed one.
Private by default. Files are auto-deleted after the retention window and are never used for advertising or to train models.

Converting scans at scale?

The same OCR pipeline is a REST API and a hosted MCP endpoint, with machine-readable discovery so scripts and agents can drive it directly.

FAQ

Common questions

Can it convert a scanned PDF to Markdown?

Yes. Image-only and scanned PDFs are OCR'd automatically into selectable Markdown – no separate OCR step and no setup. Just drop the file in the extension or web app.

Does the OCR handle other languages?

Yes. It works across many languages, including mixed-language documents, and turns the recognized text into Markdown.

The PDF has a bad text layer – can I force OCR?

Yes. Turn on force OCR so the converter re-reads the page images instead of trusting the embedded text, which fixes garbled or missing characters.

Are tables and formulas kept when converting a scan?

Yes. Scanned columns are rebuilt as real Markdown tables instead of jumbled lines, and mathematical notation is preserved rather than flattened. See extracting tables from PDF to Markdown for more.

Why is my result marked truncated?

OCR is slow, so a very long scan can hit the per-document time budget. The converter returns what it processed, flagged as a partial (truncated) result. A paid tier has a longer budget, or you can split the file.

What scan quality do I need for good OCR?

Aim for sharp, straight pages at roughly 300 DPI or higher with good contrast. Faint, low-resolution or skewed scans still convert, but accuracy drops; rescanning cleaner is the quickest fix.

Can it read handwriting?

The OCR targets printed text, so handwritten notes are not reliably recognized. Printed and typeset documents, including scans, work well.

Is it free and private?

Yes. The free tier gives 3 slots, 10 MB files, a 15-minute time budget and 1-hour retention – anonymous in the browser, no card. Files are auto-deleted after the retention window and are never used to train models.