Skip to content
PI Square

Document intelligence

Documents read as they arrive, checked against your rules, and sent where they need to go.

emailintakeinvoicecontractidclassifyread the layoutsupplierrefamountbankfields + confidencein master fileokpo existsoktotal = linesokbank matchesokvalidatebank details?check the scana person reviewserp
  1. 1Intake
  2. 2Classify
  3. 3Read the layout
  4. 4Extract with confidence
  5. 5Validate
  6. 6Post, or a person reviews
Try it

Document intelligence reads the documents your business receives, such as invoices, applications, contracts and supplier files, pulls out the facts you need and checks them against your rules. PI Square builds pipelines that pair OCR with language models, process each document as it arrives, and send anything uncertain to a person.

  1. Documents arrive

    email, uploads, scans

  2. Your cloud

    Read

    OCR and a language model

  3. Facts pulled out

    into the fields you use

  4. Checked against your rules

    anything unsure goes to a person

  5. Into your systems

    ERP, case files or a review queue

What it's for

  • Property

    Invoices and notices from many sites read and captured without retyping.

  • Procurement

    Supplier documents checked for compliance and risk before the supplier is approved.

  • Financial services

    Application packs checked for completeness and consistency before review.

  • Finance teams

    Invoices and statements matched to orders and accounts.

How we keep it safe

  • Each field carries a confidence score, and below your threshold a person checks it.
  • The original document stays linked to every record it produced.
  • Rules you can read, kept apart from the model, decide what passes.
  • Documents and the data taken from them stay in your cloud.

Built with

  • Vertex AI OCR
  • Gemini
  • Gmail and Outlook mailboxes
  • Cloud Run
  • Pub/Sub
  • Python

Where we've built this

  • A retail property business with more than 200 locations: documents that arrive by email are read and processed as they arrive.
  • A major South African retailer: supplier risk assessed from the supplier's own documents, with most of the classification automated.

How an engagement runs

  1. Understand

    We take a sample of your real documents, measure how well they can be read, and agree which fields matter.

  2. Prove

    One document type runs live, with a person checking every uncertain field.

  3. Embed

    The document workflow has a named operator, review rules and monitoring. Additional document types and changes to review thresholds are evaluated separately.

The full approach

Questions about document intelligence

What is intelligent document processing?
Software that reads documents the way a clerk would: it recognises the text, works out what kind of document it is, pulls out the facts you need and checks them. Modern systems pair OCR with a language model, so they cope with layouts they have not seen before.
What happens when the system isn't sure?
The document goes to a person with the uncertain fields marked. Their correction is recorded, and the rules and prompts are improved from those records.
Can it read scans and handwriting?
Printed scans, yes. Handwriting depends on how clear it is, so we measure it on your own samples during Understand before promising anything.
How does extracted document data connect to our ERP or accounting systems?
Extracted fields pass validation rules and are delivered directly into your systems—such as SAP, Oracle or internal databases—via secure APIs, webhooks or message queues, with the original document linked for audit.
How does the pipeline handle varying document layouts?
Modern language models interpret document structure rather than relying only on rigid coordinate templates, allowing them to extract fields across different formats. Low-confidence values are sent to human review queues.

The other five

Tell us what you're working on

Keep this general; leave out sensitive or confidential information.

Or email nikhil@pisquare.ai