EXTRACT TABLES FROM PDF

Extract PDF tables with evidence.

Review a full-table conversion or use named fields for the next step.

Rendered first page of the synthetic multi-page shipment register
Page one of a CC0 synthetic source. The outline is a review annotation, not a manual selection control.

Choose the output before you choose a workflow.

The full table and a smaller set of named fields are different deliverables. The available product routes reflect that difference.

Keep the whole table

Use the public PDF converter when the full PDF table should become a reviewable Excel workbook.

Open PDF converter

Reduce the output to named fields

Use the signed-in AI workspace when a defined set of business fields is more useful than a full-table workbook.

Open signed-in workspace

Current limitation: The current product does not offer a manual box selector.

Two source tables, one bounded workbook review.

This is a saved benchmark result, not a universal accuracy claim. The outlines document visible source tables for review. They are not product coordinates or a selectable interface.

Rendered first page of the synthetic multi-page shipment register

First page

The visible shipment table is one of two table regions represented by the saved two-page benchmark run.

Rendered continuation page of the synthetic multi-page shipment register

Continuation page

The continuation table is the second represented region. Its repeated header and rows should be reviewed with page one.

Saved run
2026-07-19, PaddleOCR-VL-1.6
Reported regions
2 table regions
Declared rows
6 of 6 matched
Declared cells
35 of 35 matched

Review the source beside the workbook.

A result can be useful without being the final record. Check every table, page boundary, header, and downstream field requirement against the original PDF.

  1. Confirm the source pages first

    Check that every page, continuation header, and source table you need is present before you compare a workbook.

  2. Review headers and repeated rows

    A repeated header or a page boundary can change how a downstream user reads the resulting workbook.

  3. Use full-table conversion for the whole table

    Open the public PDF converter when the workbook should retain the table rather than reduce it to named fields.

  4. Use named fields when the output is narrower

    Open the signed-in workspace when a reviewer needs a smaller set of defined business fields from a complex document.

What this page does not establish.

The source is a two-page selectable-text CC0 fixture. It does not establish performance for scans, handwriting, camera perspective, other languages, merged cells, naturally degraded documents, or arbitrary multi-table layouts.

The declared reference workbook is expected data for review. It is not a model output. The saved actual workbook is one dated benchmark artifact and should be compared with the source before use.

Questions before you extract a table

Can I draw a rectangle around one table?

No. The current product does not offer a manual box selector. This page documents a review path for a complex PDF, not a new selection control.

What does two table regions mean in this example?

The saved benchmark run reported table_count equal to two for this two-page synthetic source. The source has one visible shipment table on each page.

Does 35 of 35 establish a general accuracy rate?

No. It is a saved result for one dated CC0 synthetic fixture. It does not establish performance for other layouts, scans, languages, merged cells, or documents.

Which product route should I choose?

Use the public PDF converter when the full table is the output. Use the signed-in AI workspace when named, reviewable fields are the smaller deliverable.

Start with the output a reviewer can verify.

Choose the full-table converter or the signed-in field workflow, then compare each output with its source PDF.