PDF to XML

Extract PDF text items and page geometry into a documented XML structure.

Choose PDF Files

Drop files here or use the button below.

Up to 25 MB total · Files stay on your device

Table of Content10 Sections

What Does the XML Export Contain?

Export the selectable text layer of a PDF as XML. The file groups text items under their original page numbers and records page geometry, text direction and positioning transforms. It provides structured evidence for inspection and later processing.

The XML describes what the PDF text reader extracted. It is not the original XML that may have generated the document, and it does not identify invoice fields, reconstruct table relationships or export images. Image-only pages need OCR first.

PDF to XML Features

Page-Based Elements

Keep source page numbers, rotation and visible page bounds beside extracted text. This preserves a useful reference when reviewing a selected range.

Escaped Text Values

Characters such as ampersands and angle brackets remain data inside XML elements. Text resembling markup cannot become executable page content through this export.

Reversible Exceptional Strings

Characters XML cannot represent directly use an explicitly marked JSON string encoding. The marker allows a receiving program to recover those values accurately.

Browser Processing

Create the XML file on your device. The source is neither uploaded for conversion nor rewritten.

How Do I Extract XML from a PDF?

  1. Select an unencrypted PDF containing selectable text.
  2. Enter the required page numbers or ranges in Pages.
  3. Convert the file and download the XML result.
  4. Open the XML in a text editor or XML-aware application.
  5. Compare a known page and phrase before mapping elements into another system.

Choose the Source PDF

Select a PDF with a readable text layer.

Select the screenshot to enlarge it.
PDF to XML: selecting a source PDF.

Limit the Page Range

Choose the pages needed for the XML extraction.

Select the screenshot to enlarge it.
PDF to XML: page range and extraction limits.

Save the XML

Download the XML and inspect its page elements.

Select the screenshot to enlarge it.
PDF to XML: completed XML export download.

Screenshots show this toolkit using harmless sample files. Your file name, page count and settings may differ.

Why Use XML for PDF Text Inspection?

An XML-based workflow may need explicit page and item elements before applying its own transformation rules. A team can inspect those elements, retain source-page references and build a separate mapping for known document types. Keep that mapping distinct from extraction: a text item near another item does not establish a business relationship by itself.

Who Can Work with This XML Output?

Integration Developers

Use page and item elements as input to a controlled document-processing workflow.

Document Archivists

Inspect whether text remains available alongside the preserved original PDF.

Data Auditors

Compare sampled output values with their source pages before accepting transformation rules.

Settings and Output Checks

SettingWhat to Expect
SourceOne PDF up to 25 MiB, with no more than 500 source pages.
Page rangeUp to 200 selected pages. An empty range includes all pages only when within that limit.
Text limitsUp to 20,000 items on a page, 100,000 items per job and two million text characters.
DownloadAn XML document up to 64 MiB. No external DTD or schema is fetched.

Map a Known Label after Checking the Page

For a delivery document, export one page and locate a familiar reference number. Check its text transform and nearby items against the original page. Your application can then apply a rule for that document type. Avoid treating every nearby number as the same field across unrelated layouts.

Troubleshooting PDF to XML

SituationWhat to Check
The XML does not match a supplier's schemaThe export uses its own extraction structure. Create a separate mapping to the schema your receiving system expects.
An element has an encoding attributeIf encoding is json, parse its text as a JSON string to recover characters that cannot appear directly in XML.
Words are absent or ordered unexpectedlyCheck for scanned pages or unusual text placement. Apply OCR where needed and review coordinates.

FAQs

Does PDF to XML Restore the Original Source XML?

No. A PDF generally does not retain the source data model. This export records the text layer and geometry that the reader can extract.

Are PDF Tags Converted into an XML Document Model?

No. The tool does not reconstruct tagged-PDF semantics, headings or table structure. Its page and item elements describe extraction results.

Can I Convert This XML Back to the Same PDF?

The XML to PDF tool can print XML source, but it cannot recreate the original page design from this extraction file. Retain the original PDF.

XML to PDF, PDF to JSON, OCR PDF.