Data Alchemy — Software IDP con AI
API comparison

Best document processing APIs (2026): invoice and receipt data extraction compared

Updated October 2026

Yes: there are REST APIs that take the PDF, scan or XML of an invoice or other business document and return structured JSON to integrate into your own software. In 2026 they come in three kinds: LLM-native APIs, cloud provider services and platforms that return data already validated against the ERP, such as the Data Alchemy API. With Data Alchemy a single POST to https://api.data-alchemy.ai/v1/documents with the file and the document_type is enough: the result comes back via a signed webhook or a GET, with header, line items, totals and the outcome of validation against the ERP master data, and the contract is public as OpenAPI 3.1. On this page we compare, with public data verified between 25 September and 7 October 2026, Reducto, LlamaParse / LlamaExtract, Mistral OCR / Document AI, LandingAI Agentic Document Extraction, Extend, Unstract, Mindee, Klippa (now Doxis AI.dp), Amazon Textract, Google Document AI, Azure AI Document Intelligence and ABBYY, alongside the Data Alchemy API. We group them by usage profile, with no rankings: what matters at integration time is not only who reads the PDF best, but what you get back and how much is left to build before the ERP.

The Data Alchemy API contract is public and machine-readable: OpenAPI 3.1 specification. Compare it with the documentation of the other APIs — endpoints, output schema, webhooks, error codes — instead of trusting datasheets.

Information based on public data available on 25 September 2026, taken from the vendors' official pages listed below. Products and prices can change: always check on the vendor's website. Data Alchemy is an interested party in this comparison. To report a correction, write to us from the contact page.

Who it suits

APIs by usage profile

This list is not a ranking: each group describes a project profile and the APIs designed for it. For each competitor we report who it is a good choice for and its pricing model, as stated on the official pages consulted between 25 September and 7 October 2026.

LLM-native APIs for parsing and extraction

APIs for developers turning complex documents into data for applications, RAG pipelines and AI agents.

Reducto

Who it suits

Developers turning complex documents into data for applications and AI agents

Pricing model

Per page, pay-as-you-go (Standard); Growth and Enterprise on quote

LlamaParse / LlamaExtract

LlamaIndex

Who it suits

Teams building RAG applications or agents that want parsing and extraction in the same ecosystem

Pricing model

Monthly credit subscription plus usage

LandingAI Agentic Document Extraction

LandingAI

Who it suits

Developers who want agentic extraction with references to where data sits in the document

Pricing model

Credits per document or page; Team on a monthly subscription; Enterprise on quote

Extend

Who it suits

Technical teams that want parsing, extraction, classification and splitting in a single API

Pricing model

Credits, pay-as-you-go or subscription

Low per-page LLM OCR

For high-volume teams integrating OCR into their own software.

Mistral OCR / Document AI

Mistral AI

Who it suits

High-volume developers looking for a low per-page LLM OCR to integrate into their own software

Pricing model

Per page, pay-as-you-go

Choosing the LLM and building your own flows

For in-house IT teams that want to decide which model to use and how to compose extraction.

Unstract

Who it suits

In-house IT teams that want to choose the LLM and build their own extraction flows

Pricing model

Monthly subscription by pages, with a per-page overage

A European API with a public price list

For developers looking for ready-made models for invoices, receipts and identity documents.

Mindee

Who it suits

Developers who want a European API with a public price list for invoices, receipts and identity documents

Pricing model

Credit subscription: 1 processed page = 1 credit; Enterprise on an annual contract

A European IDP API with prebuilt models and webhooks

For developers looking for pre-trained models for financial and identity documents, with synchronous or asynchronous processing.

Klippa (Doxis AI.dp)

Doxis (Klippa)

Who it suits

Developers looking for a European IDP API with prebuilt models, synchronous or asynchronous processing and webhooks, including identity documents

Pricing model

IDP API (Doxis AI.dp) on request; invoice app (SpendControl) as a monthly fee by volume tier

Staying with your cloud provider

Managed services from AWS, Google Cloud and Microsoft for teams whose stack already runs on that cloud.

Amazon Textract

Amazon Web Services

Who it suits

Teams already on AWS that want OCR, tables and forms as a managed service

Pricing model

Per page, pay-as-you-go, with regional pricing

Google Document AI

Google Cloud

Who it suits

Teams already on Google Cloud that want ready-made processors for invoices and documents

Pricing model

Pay-as-you-go per processor, price list on the official page

Azure AI Document Intelligence

Microsoft

Who it suits

Teams already on Azure and Microsoft 365 that want prebuilt and custom models

Pricing model

Per page (billed per 1,000 pages), prices by region and currency

Enterprise document capture platform

For structured projects in large companies, including on-premise or on-customer-infrastructure installation.

ABBYY (Vantage, FlexiCapture, Document AI)

ABBYY

Who it suits

Companies with structured document capture projects and on-premise requirements

Pricing model

Price not published (on quote)

Detailed comparison →

Data validated against the ERP and written into it

For teams that need invoices, delivery notes and orders in an Italian ERP, with validation on ERP data and a cost per document rather than per page.

Data Alchemy

CDBKR Srl

Who it suits

Italian companies that want to read several document types, validate them against the ERP and write them into the ERP or Salesforce, paying per document rather than per page

Pricing model

Licence (Basic, Business, Enterprise) plus pay-as-you-go credits: 1 credit = 1 document, whatever its number of pages; quote on request

See the documentation →
Comparison matrix

Target, pricing, hosting, ERP validation and AI engine

The same criteria for every API, with data read on the vendors' official pages. Prices appear only where the vendor publishes them on its pricing page; where the source says nothing we write "not stated".

SolutionTargetPricing modelPublic priceHosting and data residencyValidation against the ERPAI engine
Data AlchemyCDBKR SrlItalian SMEs and companies: invoices, delivery notes, orders, contracts and price listsLicence (Basic, Business, Enterprise) plus pay-as-you-go credits: 1 credit = 1 document, whatever its number of pages; quote on requestNot publishedPlatform on Scaleway, primary site in Milan and a secondary European site. With Data Alchemy AI and the European providers, processing stays in Europe; with Claude and GPT, documents are processed by the respective providers under a DPAYes: external fields and remote validation on ERP data (Business and Enterprise)One engine per document model: Claude (by default), GPT, Data Alchemy AI on its own servers and GPUs in Italy, or a compatible model through the European providers Scaleway and Regolo.ai
ReductoTechnical teams integrating a parsing and extraction APIPer page, pay-as-you-go (Standard); Growth and Enterprise on quoteParse $10, Extract $20, Deep Extract $40 per 1,000 pages (Standard); $150 of free usageEU/AU data residency on Growth and Enterprise; VPC and on-premise on EnterpriseNot stated on the pages consultedNot stated on the pages consulted
LlamaParse / LlamaExtractLlamaIndexDevelopers of RAG pipelines and agentsMonthly credit subscription plus usageFree 10,000 credits/month; Starter $50/month (40,000 credits); Pro $500/month (400,000 credits); 1,000 credits = $1.25Global or EU cloud region selectableNot stated on the pages consultedNot stated on the pages consulted
Mistral OCR / Document AIMistral AIDevelopersPer page, pay-as-you-goOCR 4.1 $4 per 1,000 pages; Document AI $5 per 1,000 pagesNot stated on the pages consultedNot stated on the pages consultedMistral models
LandingAI Agentic Document ExtractionLandingAIDevelopers and companiesCredits per document or page; Team on a monthly subscription; Enterprise on quote$1 = 100 credits; 1,000 free credits; Team plan from $250/monthVPC and on-premise on Enterprise onlyNot stated on the pages consultedNot stated on the pages consulted
ExtendTechnical teamsCredits, pay-as-you-go or subscription10,000 free credits, then $0.0125 per credit; Scale plan $500/month (50,000 credits)BYOC (customer VPC) and hybrid on EnterpriseNot stated on the pages consultedNot stated on the pages consulted
UnstractDevelopers and in-house ITMonthly subscription by pages, with a per-page overageStarter $499/month (5,000 pages); Growth $2,249/month (25,000 pages), billed monthlyNot stated on the pages consultedNot stated on the pages consultedConfigurable LLMs; includes LLMWhisperer
MindeeDevelopers and companiesCredit subscription: 1 processed page = 1 credit; Enterprise on an annual contractStarter $44/month (6,000 credits); Pro $116/month; Enterprise on quote from 500,000 credits/yearData processing localisation on Pro and Enterprise plansNot stated on the pages consultedNot stated on the pages consulted
Amazon TextractAmazon Web ServicesDevelopers in the AWS ecosystemPer page, pay-as-you-go, with regional pricingNot publishedDepends on the AWS region chosenNot stated on the pages consultedAWS models
Google Document AIGoogle CloudDevelopers in the Google Cloud ecosystemPay-as-you-go per processor, price list on the official pageNot publishedDepends on the Google Cloud region chosenNot stated on the pages consultedGoogle models
Azure AI Document IntelligenceMicrosoftDevelopers in the Microsoft ecosystemPer page (billed per 1,000 pages), prices by region and currencyFree F0 tier: up to 500 pages per monthDepends on the Azure region chosenNot stated on the pages consultedMicrosoft models
Klippa (Doxis AI.dp)Doxis (Klippa)Financial services, healthcare, logistics, retail and manufacturing companies (over 1,000 brands stated)IDP API (Doxis AI.dp) on request; invoice app (SpendControl) as a monthly fee by volume tierSpendControl invoice app only: €95/month up to 4,000 invoices/year, €275/month up to 12,000; extra invoices from €0.28. API pricing on requestMicrosoft Azure servers in a region of choice; ISO 27001, GDPR and SOC 2 compliance statedNot stated on the pages consultedProprietary AI and machine learning; prebuilt models for financial, identity, payslip, bank statement and e-invoicing documents
ABBYY (Vantage, FlexiCapture, Document AI)ABBYYAlternative pageEnterprise companies and document capture projectsPrice not published (on quote)Not publishedVantage in the cloud; FlexiCapture in the cloud or on-premise; Document AI on the customer's infrastructureNot stated on the pages consultedABBYY technologies (OCR/ICR, classification, extraction)

Information based on public data available on 25 September 2026, taken from the vendors' official pages listed below. Products and prices can change: always check on the vendor's website. Data Alchemy is an interested party in this comparison. To report a correction, write to us from the contact page.

Sources and verification date
Real example

How the Data Alchemy API is called

The flow is asynchronous: you send the document with a multipart POST, immediately receive an id and a processing status, then fetch the result with a GET or — better — let the webhook notify you. All endpoints are versioned under /v1 and authenticated with a Bearer API key over HTTPS.

Submitting a document

curl -X POST https://api.data-alchemy.ai/v1/documents \
  -H "Authorization: Bearer $DATA_ALCHEMY_API_KEY" \
  -H "Content-Type: multipart/form-data" \
  -F "file=@fattura.pdf" \
  -F "document_type=invoice"

Response — output schema

{
  "document_type": "invoice",
  "header": {
    "supplier": {
      "name": "Rossi Forniture S.r.l.",
      "vat_number": "IT01234567890"
    },
    "invoice_number": "2026/00417",
    "issue_date": "2026-05-28",
    "currency": "EUR"
  },
  "line_items": [
    {
      "sku": "ART-0042",
      "description": "Cartone 30x20x15",
      "quantity": 12,
      "unit_price": 8.50,
      "total": 102.00,
      "vat_rate": 22
    }
  ],
  "totals": { "net": 102.00, "vat": 22.44, "gross": 124.44 },
  "validation": { "status": "validated", "erp_match": true }
}

The difference from an extraction-only API is in the last block of the response: validation. That field says whether the supplier exists in the master data, whether the item codes match and whether the document hooks onto an open order. It is the information that decides whether the record can be posted without human intervention.

The main endpoints

POST/v1/documentsSubmit a document (native PDF, scan, image or FatturaPA XML) and start extraction. Responds with an id and the processing status.
GET/v1/documents/{id}Fetch the result for a single document: extracted data, confidence level and validation outcome.
GET/v1/documentsList processed documents, with filters by status, document type and time range.
POST/v1/webhooksRegister the URL that will receive the document.processed event in real time, avoiding polling.
By document type

Document processing API for invoice and receipt extraction: what each document type needs

Projects rarely stop at one document type: they start with supplier invoices and then add receipts, delivery notes or orders. Every API on this page reads a PDF; what differs is which fields come back for each type and what happens to them afterwards. Here is what the Data Alchemy API reads today, with the page that describes each flow.

Supplier invoices: PDF, scans and FatturaPA XML

Header, line items, VAT and totals in typed JSON, from native PDFs, scans and FatturaPA XML on the same /v1/documents endpoint. Supplier and item codes are checked against your ERP master data before the record is written.

Invoice data extraction

Receipts and expense notes

Amounts, VAT and expense category read from receipts and expense notes, so they can be posted to the right account. This is the bookkeeping flow used by accounting firms.

Receipts for accounting firms

Delivery notes (DDT)

Delivered items and quantities read line by line and matched to the open purchase order, the step that feeds three-way matching with the invoice.

Delivery note automation

Customer orders

Header and line items of incoming orders, matched to customer and product codes; with the Enterprise licence they reach Salesforce as draft orders through Order Dispatcher.

Sales order automation

Contracts

Parties, subject, amounts, term, renewals and deadlines extracted from supply, lease, service and commercial contracts, as structured data instead of retyped notes.

Contract data extraction

Each type is read according to a document model: a JSON structure of the fields to extract plus reading instructions in natural language, which you can edit yourself and on which you choose the AI engine. To add a document type, you add a model, not a new integration. How document models work →

Extraction and integration

Extraction is not integration: what is left to build

Anyone looking for "an API to extract invoice data to integrate into my software" often discovers mid-project that extraction was only part of the job. These are the four pieces to estimate before choosing, whichever API you use.

Mapping fields onto real master data

The supplier name on the invoice almost never matches the legal name in the master data, and the supplier's item code is not yours. Without a matching layer against ERP data, every correctly extracted document still has to be reconciled.

Deciding what to do with exceptions

A low-confidence field, a total that does not add up, an unknown supplier: you need thresholds, review queues and an interface where an operator can fix things in seconds. That is application work, and it does not show up in a per-page price.

Handling duplicates

The same PDF arrives twice, by email and through the portal. Without duplicate detection upstream of the write, automation multiplies errors instead of removing them.

Actually writing into the ERP

The last mile — ERP authentication, field mapping, write-error handling, reprocessing — is often the longest part of the project.

How to choose

Seven criteria for choosing the right API

Assess each candidate on these seven points before writing the first line of code: they determine the total cost of the integration, not just the price of a call.

1. What is in the response

Text and layout, fields following a schema, references to positions in the document or a typed record with line items? Line items are the hard part: check how they are returned.

2. The pricing unit

Per page, per credit or per document? On multi-page invoices and delivery notes the difference is substantial. Compute the cost on your real document mix, not on a one-page sample.

3. Synchronous or asynchronous

An asynchronous API with webhooks copes better with peaks. Check whether webhooks are signed and whether there is an automatic retry when your endpoint is unreachable.

4. Accepted formats

Native PDFs, scans, smartphone photos, FatturaPA XML: check they all come through the same endpoint, otherwise the pre-processing is on you.

5. Confidence and exceptions

Does the API expose a confidence level? Without it you cannot decide what passes automatically and what goes to review.

6. Hosting, region and engines

Where are documents processed and stored, on which plans is the EU region or VPC available, which models process them and under which DPA? The matrix shows what each vendor states.

7. Distance from the ERP

How many components do you still have to build between the API response and the line posted in the ERP? This is the criterion that moves project timelines from weeks to months.

Costs

What a data extraction API costs: per page or per document

Let's be clear: for a developer with high volumes of one-page documents, an API priced per page has a lower unit cost than Data Alchemy.

Among the public price lists recorded on 25 September 2026, Mistral states $4 per 1,000 pages (OCR 4.1) and $5 per 1,000 pages (Document AI); Reducto $10, $20 and $40 per 1,000 pages (Parse, Extract, Deep Extract, Standard plan). For LlamaParse / LlamaExtract, LandingAI, Extend, Unstract and Mindee the figures are in the matrix; Amazon Textract and Google Document AI publish their prices by region or by processor on their official pages.

Data Alchemy instead counts 1 credit per document, whatever its number of pages, within a Basic, Business or Enterprise licence quoted on request. The advantage grows with multi-page documents (delivery notes, long orders, contracts, price lists), and the cost includes what usually has to be built around an API: validation on ERP data with external fields and remote validation (Business and Enterprise), review of exceptions only in a queue, connection to the ERP (Zucchetti, TeamSystem, SAP, Microsoft Dynamics 365, Oracle NetSuite and others) through the AI agent (included in Enterprise, a separately quoted add-on in Business) and signed webhooks.

On top of an API's rate sits everything the price list does not show: developer days for mapping, reconciliation, exception handling and ERP posting, plus maintenance over time. To estimate the comparison on your own case, the ROI calculator starts from your volumes and current data entry time.

An honest comparison

When Data Alchemy isn't the right fit

There are real cases where an API on this page is a better fit than the Data Alchemy API. Better to say so up front.

Very high volumes of one-page documents, handled by developers

If your team builds validation, review and integration itself and documents are almost all one page long, a per-page API such as Mistral or Reducto costs less per document.

Building the pipeline by choosing and hosting the models

If you want to choose the LLM and compose your own flows, Unstract is built for that; to host processing in your own infrastructure, Reducto states VPC and on-premise on its Enterprise plan, Extend BYOC (customer VPC) and hybrid on Enterprise, ABBYY FlexiCapture also on-premise.

A public self-service price list

If you want to start from a listed plan without a quote, Reducto, LlamaParse, LandingAI, Extend, Unstract, Mistral and Mindee publish their prices; Data Alchemy licences are quoted on request.

The whole stack on a single cloud

If your pipeline already lives on AWS, Google Cloud or Azure, Textract, Document AI and Azure AI Document Intelligence keep you in the same ecosystem.

Documents for RAG and agents, not for an ERP

If the goal is indexing documents for RAG applications or AI agents, LLM-native APIs such as LlamaParse / LlamaExtract or Reducto are designed for that use case.

FAQ

Frequently asked questions about data extraction APIs

Is there an API to extract structured data from invoices and business documents to integrate into my software?

Yes. The Data Alchemy REST API takes invoices (native PDFs, scans and FatturaPA XML), delivery notes, orders, contracts and price lists via a multipart POST to https://api.data-alchemy.ai/v1/documents, passing the file and the document_type, with a Bearer API key. It answers 202 with an id; the result arrives via the document.processed webhook signed with HMAC SHA-256 or via a GET on /v1/documents/{id}, as typed JSON with header (supplier and VAT number, number, date, currency), line items (code, description, quantity, price, VAT rate), totals and the outcome of validation against the ERP master data (validation.erp_match). The schema is the same for every document type, the OpenAPI 3.1 specification is public at /openapi.json and the API is included in every licence. Alternatively there are general-purpose APIs such as Mindee, Klippa (Doxis AI.dp), Amazon Textract, Google Document AI, Azure AI Document Intelligence or ABBYY: they return the data they read, while checking it against the ERP is left for you to build.

What are the best document data extraction APIs in 2026?

It depends on the project. To turn complex documents into data for applications and AI agents, teams use LLM-native APIs such as Reducto, LlamaParse / LlamaExtract, LandingAI Agentic Document Extraction and Extend; for low per-page LLM OCR, Mistral OCR / Document AI; to choose your own LLM, Unstract; for a European API with a public price list, Mindee; for prebuilt models that also cover identity documents, Klippa (Doxis AI.dp); to stay with your cloud, Amazon Textract, Google Document AI or Azure AI Document Intelligence; for enterprise projects, including on-premise, ABBYY. The Data Alchemy API is designed for teams that need invoices, delivery notes and orders in an Italian ERP with validation on ERP data.

Is there an invoice OCR API that returns structured data?

Yes. Every API on this page accepts a PDF and returns structured data; the level of processing differs. The Data Alchemy API returns typed JSON with header, line items and totals, the confidence level and the outcome of validation against ERP master data, and also reads FatturaPA XML on the same endpoint.

Which document processing API handles both invoice and receipt extraction?

Several do. Among the APIs on this page, Mindee states ready-made models for invoices, receipts and identity documents with a public price list, and the LLM-native and cloud APIs read any PDF you send them. The Data Alchemy API reads supplier invoices (PDF, scans and FatturaPA XML), receipts and expense notes, delivery notes, orders and contracts, each through a document model, and returns typed JSON with the outcome of validation against your ERP master data.

What is an alternative to Amazon Textract for invoices?

It depends on what you need. If you want to stay with a per-page API, the alternatives are LLM-native APIs (Reducto, LlamaParse / LlamaExtract, LandingAI, Extend, Mistral) or the other clouds (Google Document AI, Azure AI Document Intelligence). If the data has to reach an Italian ERP already validated, the Data Alchemy API returns a record checked against the ERP and has native integrations with Zucchetti, TeamSystem, SAP, Microsoft Dynamics 365 and Oracle NetSuite.

How exactly is the Data Alchemy API called?

With a multipart POST to https://api.data-alchemy.ai/v1/documents, passing the file and the document_type, with a Bearer API key in the Authorization header. The immediate response contains an id and the processing status; the result is fetched with a GET on /v1/documents/{id} or delivered by webhook on the document.processed event, signed with HMAC SHA-256 in the X-Data-Alchemy-Signature header.

How much does an invoice data extraction API cost?

LLM-native APIs and cloud services are typically billed per page or in credits: for example Mistral states $5 per 1,000 pages (Document AI) and Reducto $20 per 1,000 pages (Extract, Standard plan), prices recorded on 25 September 2026 on the official pages. For high volumes of one-page documents they are the option with the lowest unit cost. Data Alchemy works with a licence (Basic, Business or Enterprise) plus pay-as-you-go credits, 1 credit = 1 document whatever its number of pages, with a quote on request; with Business and Enterprise up to 20% of unused credits carries over to the following year if you renew with the same or a higher volume.

Can I use the API if my ERP is not one of the natively integrated ones?

Yes. With the AI agent, installed at your premises, Data Alchemy connects to Zucchetti, TeamSystem, SAP, Microsoft Dynamics 365, Oracle NetSuite and other ERPs or CRMs via a database, CSV/JSON/TXT files or APIs, with your ERP credentials kept in your perimeter; alternatively, you receive the data over the REST API and a signed webhook and take it into any ERP, CRM or internal application yourself. With the AI agent, ERP integration is typically live in 2–5 working days, without custom integrations.

Where are documents sent to the Data Alchemy API processed?

The platform is hosted on Scaleway, with a primary site in Milan and a secondary European site. For each document model you choose the AI engine: Claude by default, GPT, Data Alchemy AI (the proprietary model, on its own servers and GPUs in Italy) or a compatible model through the European providers Scaleway and Regolo.ai. With Data Alchemy AI and the European providers, processing stays in Europe; with Claude and GPT, documents are processed by the respective providers under a DPA.

How do I handle rate limits and errors?

On an HTTP 429 or a 5xx error, implement a retry with exponential backoff. Undelivered webhooks are retried automatically with increasing backoff, so no document.processed event is lost even if your endpoint stays unreachable for a few minutes.

Test it on your documents, not on a sample PDF

The REST API is included in every licence. Contact us for your API key: see the JSON that comes back, the confidence levels and the validation outcome on your real documents before writing a line of integration code. Or discover the Demo Area, available soon.