OCR-Genius - Back to home
ExtractAn OCR-Genius product

From logistics document to normalized data your system already understands.

Extract processes invoices, packing lists, declarations and transport forms. It turns items that arrive in different formats into a predictable schema — yours — ready for ERP, customs or warehouse.

See how it works
  • Built for logistics
  • Schema made to measure
  • Webhook · Excel · API
  • +97% accuracy

Extract pipeline — from document to data in minutes

01 · IntakeSource documentInvoices, packing lists, DUCA, B/L. PDF, image or API.
02 · ExtractionExtraction and disambiguationAgentic AI that reads the right field, not the position.
03 · NormalizationMatch to your fieldsEvery item lands in your schema, validated per regulation.
04 · OutputData ready to integrateWebhook, Excel or dashboard — straight into your ERP, WMS or TMS.

From document to data in minutes — +97% accuracy, no manual intervention.

The Extract pipeline in four steps: the source document comes in, its fields are extracted and disambiguated, matched against your schema, and the data comes out ready to integrate via webhook, Excel or dashboard.
Day to day in operations / customs

Every supplier sends the same kind of document… in a different shape

A single clearance can require hundreds of lines copied from PDFs, scans or photos. One wrong digit and the shipment stops.

How it arrives today
Scanned packing list with inconsistent columns and descriptions
  • Mixed fields
  • Unformatted units
  • Duplicate keys
What your ERP needs
Example of a packing list normalized to the customer's schema
skuproductocolororigenpeso_kg
3MF10730622Sports shoeWhiteVN1.880
3MF10263563Sports shoeCreamVN17.360
3MF10070692Athletic footwearMidnightID1.640
  • One field, one value
  • Your system's keys
  • +97% accuracy
01

Manual work

Hours typing invoices, delivery notes and packing lists.

02

Unpredictable formats

Different columns per issuer and mixed-up fields.

03

Generic OCR is not enough

Plain text with no normalization layer.

04

Operational pressure

More volume, more regulation, thinner margins.

Section wrap-up

You don't need “more OCR”. You need predictable items carrying your system's keys.

The Extract method

Four steps. From file to usable data.

OCR-Genius understands the logistics document. Extract normalizes it to your schema and delivers it where your stack runs.

One PDF with four documents inside

Auto-split — classification by type before extraction
01 · Intake
input.pdf1 file · 12 pages4 documents inside, undeclared
Auto-split — classification by type
  • Commercial invoiceInvoice · pp. 1–3
  • Packing listPacking List · pp. 4–6
  • Certificate of originEUR.1 · p. 7
  • Bill of ladingBill of Lading · pp. 8–12
02 · Extraction
ExtractionEach type, with its own rules.

Auto-split: if you don't declare the type, Extract detects it, separates it and routes each document on its own.

A 12-page PDF is automatically split into invoice, packing list, EUR.1 certificate and bill of lading, each routed to extraction
1

Intake

Native or scanned PDF, image or API. Auto-split if you don't declare the type.

2

Extraction

Headers and items. Disambiguates brand, origin, composition.

3

Normalization

Match to the fields you defined — with your names.

4

Delivery

Signed webhook, Excel or dashboard. Versioned template.

The promise

Upload the document. Get back the JSON, Excel or webhook your system expects

Tailored to you, not to us. Extract does not sell “generic OCR fields”: it delivers normalized items with the keys, the order and the rules of your operation.

See how it works
Upload → payload (typical)
<60 sUpload → payload (typical)
Match with your schema
100%Match with your schema
Logistics accuracy
+97%Logistics accuracy
Dashboard · API · webhook
3Dashboard · API · webhook
The real value

What arrives from the supplier is not what your system consumes

Extract closes that gap without an operator rewriting every line.

Suppliers
ADescription · Qty · Price
BDesc. · Cantidad · Precio
CProduct Description · quantity · amount
ExtractExtracts · disambiguates · maps to your fields
Your schema
Descripciontext
Cantidaddecimal
Preciodecimal
Three suppliers with different field names are mapped, through Extract, to the Descripcion, Cantidad and Precio fields of your schema
Before — what arrives

Loose columns, unpredictable names

product_description_[text]
BREWGILL 45/82 EXP GB170
shipped_quantity
24,000.20
Price
0.84
After — your company

Fixed rows, your keys

Descripcion
BREWGILL 45/82 EXP GB170
Cantidad
24000.20
Precio
0.84
Stage by stage

What happens to every document

Stage 1 · Intake

It arrives through the channel you already use

A web panel for operations, a REST API with a per-organization key, or email to an address of your own. Leave out the document type and Extract detects it and splits it inside the same PDF.

Option 1Web panel
Drag files onto the list
Upload folderUpload document
invoice_3421.pdfinvoice

For the operations team — no integration needed.

Option 2REST API
curl -X POST \https://…ocr-genius.com/api/v1/ingest \-H "Authorization: Bearer ocrg_…" \-F "file=@/path/to/invoice.pdf" \-F "document_type=packing_list"
multipart/form-data · per-organization key

To integrate your ERP, TMS or WMS — or to send in batches.

Option 3Email
docs@your-org.ocr-genius.com
Attachment — packing_list.pdf
Attachment — bl_MSC4471.pdf

Forward what you already receive — attachments come in on their own.

POST /api/v1/ingestLeave out the document type and Extract detects it and splits it inside the same PDF.
202 · job_id returned
Three intake channels — web panel, REST API and email — converge on POST /api/v1/ingest.
Stage 2 · Extraction

Everything on the page, disambiguated

Header and line items, row by row. What comes apart from the description —brand, origin, composition, size— is not lost: it stays on the original items and is promoted to a field if you asked for it.

DescripcionCantidadCodigoColorTotalLineaCountryCodeNCM
Everything on the page, disambiguated
Stage 3 · Normalization

The schema is yours, not ours

You define your fields with a name and a type. Mapping lines up the columns found in the document with your destinations, and you drag it if you want to change it: the preview updates live and saving re-projects the delivery.

7 mapped1 unmappedversioned template
The schema is yours, not ours
Stage 4 · Delivery and control

Webhook, Excel or panel — with an audit trail

Every document leaves a record of events. The template is versioned: you remap, reprocess, restore an earlier version or resend the webhook. No more “the OCR failed, upload everything again” loop.

OutputThree ways to receive it
Signed webhookPOST https://your-erp/hooks/ocr · HMAC
Excel / CSVdelivery_3421.xlsx · your columns
Web panelReview and download per operation

Same payload, three channels — you pick by use case.

EventsEnd-to-end trail
14:27:02document.received
14:27:19extraction.completed
14:27:24mapping.applied
14:27:26delivery.sent · 200

Every document leaves a record — no black boxes.

ControlYou fix things without starting over
RemapChange the mapping without uploading again
ReprocessRetries extraction on the same file
ResendFires the webhook at the current endpoint again

No more “the OCR failed — upload everything again” loop.

template v3 · activeEvery mapping change creates a version. Restore an earlier one and the delivery is re-projected.
restore v2
Delivery goes out by signed webhook, Excel or panel; every document leaves an event trail and the template is versioned.
Documents

Built for the paperwork of the supply chain

Forms with line items and header-only documents. Each type is enabled per organization, with its own fields and template.

  • inv

    Commercial invoices

    Goods lines with prices, weights and codes.

  • pl

    Packing lists

    Containers, packages and cargo detail.

  • dua

    DUA / DUCA / DUM

    Customs declarations with line items, in their Spanish, Central American and North African variants.

  • eur

    EUR.1

    Certificates of origin.

  • bl

    Bill of Lading

    Bills of lading — header and cargo detail.

  • +

    Other forms

    New types by configuration, not by rewriting code.

Differentiators

Why Extract and not “another OCR”

Six product decisions you notice on day one of operation.

01 — 06
  1. The output has the shape of your system

    JSON, Excel or webhook with your fields — no mapping layer afterwards.

  2. Predictable line items, not loose text

    Typed rows out of a different layout from every supplier.

  3. Three ways in, one funnel

    Panel, API and email land in the same pipeline and the same template.

  4. Built for the operator

    Remap, reprocess, restore, resend — with an audit trail per document.

  5. Test before you publish

    You see the payload that would have been sent. No surprises in production.

  6. The OCR-Genius family

    The same logistics document AI; Extract normalizes and delivers.

Integrations

It comes in through the channel you already use. It leaves in the shape your stack expects.

We are the layer between the supplier's document and your ERP, TMS or WMS.

Input
Web panel
REST API
Email · batch
Extract pipelineconfig · extraction · template
Output
Signed webhook
Excel / CSV
Web panel
Web panel, REST API and email feed the Extract pipeline, which delivers a signed webhook, Excel or the panel.
  • api

    REST ingest

    POST /api/v1/ingest

  • hook

    HMAC webhooks

    Payload in your template, signed.

  • xls

    Excel export

    Headers, line items or both on one sheet.

  • erp

    ERP · WMS · TMS

    By webhook or scheduled export.

“We already have an OCR service”

That gives you text or generic fields. Extract delivers the exact payload your ERP expects.

Fit

Built for people who live off cross-border trade documents

If you recognize yourself in the ideal profile, rollout is measured in days.

Ideal profile
  • Customs brokers and clearing agents
  • Importers and exporters with more than 50 supplier documents a week
  • Freight forwarders and multi-country 3PL / 4PL operators
  • Teams with an ERP, TMS or WMS who do not want to rewrite mappings
  • Too much volume for Excel; too vertical for a generic OCR
Not your moment yet
  • Fewer than a dozen documents a month
  • A single supplier with a stable format, already integrated over EDI
  • Documents with no line items and no schema of your own, where reading the text is enough

Write to us anyway: if the volume is going to grow, it pays to design the schema before the debt piles up.

Trust

Multi-tenant

Your documents are sensitive commercial data. The platform is built accordingly.

multi-tenanthmacgdpr
  • Isolation per organization

    Data, configuration and templates kept separate per tenant.

  • Users and roles

    Authentication per organization, with per-user permissions.

  • Hashed API keys

    Revocable keys, never stored in plain text.

  • Signed webhooks

    HMAC signature to verify the origin of every delivery.

  • Encrypted in transit and at rest

    Cloud infrastructure in certified data centers.

Traceability
  • Audit trail for configuration, deliveries and resends
  • Template version history, with restore
  • A record per document: who, when and with which template
Compliance

Data processing aligned with the GDPR

We send you the technical detail and the processing agreement before the trial.

How to start

From the first test invoice to the webhook in production

Five steps. We do the first two with you.

  1. Samples

    You send us your three worst PDFs.

  2. Organization

    We create your space and enable the document types.

  3. Fields

    You define your schema and publish it.

  4. Template

    You pick the output: JSON, Excel or webhook.

  5. In production

    API key and destination URL. Off you go.

The trial

Send us 3 of your worst invoices / packing list / any other document. Within 24 hours we send back the normalized JSON / excel.

Send my documents
Frequently asked questions

What people usually ask us before the trial

Start processing your documents with efficiency and accuracy

Transformation from disorganized documents to structured data