SmartLink
AI & automation

AI document processing in Pakistan: invoices, orders and forms

AI document processing in Pakistan is the most measurable thing an enterprise can buy from this field, because the work it replaces is already counted in hours. SmartLink Services builds extraction pipelines from Karachi on Azure AI Document Intelligence and language model APIs, validating every field against ERP master data and routing only the uncertain cases to a person. What follows is the commercial detail. Why nobody can quote a price from a brochure, what a pilot on your own documents costs and reveals, how long it runs, and the volume below which this is not worth automating at all. We scope a pilot on one document type before quoting a build.

Inputs
PDF \u00b7 scan \u00b7 email
Validation
Against ERP master data
Human
Exception review only
Overview

Extraction is the easy part. The exception path is the work.

Reading fields off an invoice is close to a solved problem. Modern extraction handles a supplier name, a total, a tax number and a set of line items on most documents most of the time. What is not solved, and what decides whether an automation project succeeds, is what happens on the documents where it is wrong or uncertain. Blind automation posts those errors into your ledger, where they cost far more to find than they would have cost to check.

So the design starts from the exception path. Every field gets a confidence signal and a validation rule. Documents that pass both go through automatically. Documents that fail either go to a review queue, where a person sees the extracted value next to the original image, corrects what is wrong and approves. The measure of the system is not extraction accuracy in isolation. It is how much work goes through untouched, and how quickly a reviewer can clear what does not.

Before any of that, we want a definition of done: which document types, which fields, what counts as correct, and a set of real documents labelled by somebody who knows the answer. Without that there is no way to tell whether a change has improved anything. With it, model selection becomes an evidence based decision rather than a preference, and later changes can be tested rather than hoped about.

Scope

What AI document processing covers

Everything below is agreed in writing before any ai solutions work starts, so both sides know what is in and what is not.

What the work covers

  • Document type analysis and field extraction design
  • Model selection or fine-tuning against your document samples
  • Validation rules checked against ERP master data
  • Confidence thresholds and exception queue design
  • Posting into ERP with full audit trail
  • Accuracy monitoring and periodic retraining

What you get at handover

  • Deployed extraction pipeline
  • Accuracy benchmark by document type
  • Exception review interface
  • Audit trail of automated postings
  • Retraining and monitoring procedure

Typically involves

ERP APIs
Discuss this service
01

Define correct before you choose anything

Document projects that start with a product demonstration tend to end in argument, because nobody agreed what accuracy means. Is a supplier name correct if it reads Limited where the document says Ltd. Is a total correct if the currency is missing. Does a missing purchase order reference count as an extraction failure or a document failure. These sound pedantic and they are exactly the questions that decide whether a pilot is judged a success or a disappointment.

We build a labelled set first, taken from your real documents rather than samples, covering the ordinary cases and a deliberate share of the awkward ones: the poor scan, the handwritten annotation, the two page invoice, the credit note, the supplier who changed their template last year. That set becomes the benchmark, and every subsequent change is measured against it. It also makes vendor claims testable, since a general accuracy figure means nothing against your particular mix of documents.

Definition of done is a business decision as much as a technical one. For accounts payable it might be that a stated share of invoices post without human touch while the rest are cleared within an agreed time. That is a target you can measure and argue about honestly. What we avoid is a project whose success criterion is that the system is live, because that criterion is met perfectly well by a system nobody uses.

  • A labelled benchmark set built from your own documents, including the awkward ones
  • Field level definitions of correct, agreed with the people who do the work today
  • Success stated as a measurable outcome rather than as the system being live
  • Vendor and model claims tested against your document mix, not against general figures
  • The benchmark retained, so later changes can be measured rather than assumed
02

Confidence thresholds and the queue a person actually works

Confidence is only useful if it is calibrated and acted upon. A pipeline reporting high confidence on everything, including the fields it got wrong, is worse than one with no confidence signal at all, because it invites trust it has not earned. We check confidence against the labelled set to see whether low scores genuinely correlate with errors, then set thresholds per field, since a total needs more certainty than a delivery note reference.

The review interface is where most of the value is won or lost. A reviewer needs the document image and the extracted values on one screen, with the field highlighted in the source, keyboard navigation, and the ability to correct and approve without reaching for a mouse. If reviewing takes longer than the original manual entry, people will stop using the system and go back to typing, and no amount of extraction quality will rescue it.

Corrections should feed back. Every correction is a labelled example, and collecting them gives you both an ongoing accuracy measure and material for improvement. What we do not do is retrain automatically on corrections without review, because that is how a systematic reviewer error becomes a systematic model error. Retraining is a deliberate act with a benchmark run before and after, not a background process nobody is watching.

  • Confidence calibrated against the labelled set rather than trusted as reported
  • Thresholds set per field, since a total warrants more certainty than a reference number
  • A review screen showing document and extracted values together, operable from the keyboard
  • Corrections captured as labelled examples and used as an ongoing accuracy measure
  • Retraining performed deliberately, with a benchmark run before and after, never silently
01

What AI document processing costs in Pakistan

Price follows the task here, and any supplier quoting from a menu is quoting for a demonstration. Reading one supplier's invoice layout into an ERP and reading fourteen layouts including two handwritten delivery notes are not the same job, do not need the same people, and cannot come from the same table. What can be compared are rates and the shape of the engagement, so ask for both.

Rate cards in this market are public enough to be useful. Junior work is quoted at PKR 500 to 1,000 an hour, full stack development at PKR 3,500 to 8,000, and senior specialists at the top of that band. The nearest published project comparator is custom web application work at PKR 200,000 to 1,000,000 and upwards, which is fair because an extraction pipeline is software wrapped around a model rather than a model on its own.

Four things move the number. Document variety is the largest, since every supplier who redesigns an invoice is another case to handle and a scanned fax is not a clean PDF. Whether anybody has ever recorded the correct answers comes second, because building that set is real work with no shortcut and it is where the honest projects spend their first fortnight. Integration is third: posting into an ERP with a three way match against a purchase order and a goods receipt is considerably more than producing a spreadsheet. The exception route is fourth, and it is the part most often missing from a proposal, because somebody needs a screen showing what the system was unsure about and why.

Two costs sit outside the build and are the ones clients meet later. Model consumption is metered by the provider and rises quietly as more teams find the tool useful. Exception review is staffed by people who need more expertise than the team that previously typed everything. We scope and price a pilot on one document type first, and quote the build once that pilot has produced a measured number. A real figure follows discovery.

  • Rate cards comparable across suppliers, since a menu price for extraction is a demonstration price
  • Custom web application pricing at PKR 200,000 to 1,000,000 and upwards as the nearest comparator
  • Document variety counted by layout rather than by volume, because layouts drive the effort
  • Recording correct answers on real documents treated as funded project work
  • Metered model consumption and reviewer time budgeted separately from the build
02

How long a document processing pilot takes

Reported delivery timelines in this market give six to twelve weeks for a focused single scope go live, and a pilot on one document type fits that band. Six weeks is realistic where the documents are digital, the fields are agreed and a person can be freed to confirm correct answers. Twelve is realistic where scans are involved, several suppliers use their own layouts, or the ERP write has to pass a three way match.

Unlabelled documents are the commonest reason a plan slips. The invoices exist, the delivery notes exist, and nowhere has anybody written down what the right answer was for each field. Assembling a few hundred real cases with agreed correct answers takes days of an experienced person's time and surfaces disagreements between two reviewers who both handle these documents every week. That disagreement is useful information about your process, and it is not fast to resolve.

Master data is the quiet one. Validation only works where a supplier can be matched, and a supplier master carrying three spellings of one company will send perfectly good extractions to the exception queue. Cleaning that is often the highest value week of the whole project, and it improves every report you already run.

Access has its own timetable. An extract of a year of documents needs a named owner, a window and sometimes a legal view on what may leave your environment. Those go on the plan as dependencies with dates rather than being discovered in week four, by which point the pilot has a deadline and no data.

  • Six to twelve weeks reported for a focused single scope go live in this market
  • A few hundred real documents with agreed correct answers assembled before anything is judged
  • Supplier and item master data cleaned early, since validation depends on it entirely
  • Document extracts scheduled with an owner, a window and a view on what may leave
  • Scans, handwriting and multiple supplier layouts recognised as separate cases
03

When the volume does not justify automation

Do the arithmetic before anybody builds anything, because it settles the question in an afternoon and it frequently settles it against us. Count the documents a month. Time how long one takes a competent person to key, including the checking. That gives the hours the automation could return. Then set them against what the work costs to build and what it costs to run, which is metered model consumption plus the reviewer who works the exception queue every morning.

Two hundred invoices a month, at two minutes each, is under seven hours of work. No extraction pipeline pays that back, and a supplier who tells you otherwise is selling a project rather than an outcome. Four thousand invoices a month across sixty suppliers is an entirely different conversation, and the case usually makes itself without anybody having to be persuaded of anything.

Volume is not the only test, and two others matter more than they look. Verification comes first: if nobody can tell whether an extracted value was right, the queue will fill with things nobody can resolve and the automation quietly becomes a second data entry job. Consequence comes second. Where an error posts a payment to the wrong supplier, the review that would catch it has to be funded, and if it cannot be funded then the honest answer is that this is not ready to automate.

There is usually a cheaper answer sitting in plain sight and we would rather name it. Ask your ten largest suppliers for a structured file instead of a PDF, which many will provide because it saves them work too. Replace a form somebody retypes with a portal the customer fills in directly, and you have deleted the extraction rather than automated it. Fix a supplier master so a lookup succeeds. Each of those costs less than a model, runs without a monthly bill and cannot be wrong about a decimal point. Where they cover the requirement, that is what we will quote for, and we will say so before you have spent anything with us.

  • Documents a month, minutes each and reviewer cost calculated before any build is scoped
  • Verification tested first, since an answer nobody can check cannot be improved
  • High consequence postings automated only where the review that catches errors is funded
  • A structured file requested from large suppliers as the cheaper alternative to extraction
  • A portal or corrected form recommended where it deletes the document entirely
How we deliver

Delivering AI document processing

One process, a cost in hours today, and a target that can be measured afterwards. These six steps add something the other practices do not need: an evaluation set built before anybody writes a prompt.

  1. 01

    Discover

    Volume, consistency, how an answer would be checked, and what a wrong one costs. We also give a capable person the same documents as a control, because if they cannot produce the answer, no model will.

  2. 02

    Blueprint

    A few hundred real cases with agreed correct answers, assembled with the people who do the work now. Where two experienced reviewers disagree, that is an undocumented rule, and it gets settled here.

  3. 03

    Build

    Extraction, validation against the purchase order or master record, confidence thresholds and the exception queue are built together. Retrieval is filtered by the asking user's existing permissions, never after the answer is generated.

  4. 04

    Test

    The evaluation set is run and scored against the human baseline, per document type, so you can see where accuracy holds and where it does not. Thresholds are set from that, not from a default.

  5. 05

    Go live

    Live on one process, with cost budgets and alerts already configured, exception owners named, and a written route back to the manual process that the people who would use it have rehearsed.

  6. 06

    Run

    Re-run the evaluation set on a schedule and after any model or prompt change, because a supplier redesigning an invoice can degrade accuracy quietly. Below the agreed threshold, the automation pauses.

Working together

Where this is not worth automating

Low volume, high variety document types are usually not worth the effort. If a category produces a few documents a month and every one is different, the build cost, the review interface and the ongoing monitoring will exceed the time saved, and we will say so rather than take the work. The same applies where the underlying process should change: a form that could be a web form, a report that could be a data feed, an approval that exists only because it always has.

The cases that repay the investment are steady, repetitive and validated against data you already hold. Supplier invoices, purchase orders, delivery notes and standard forms have all three properties. There the automation is measurable, the exception path is manageable, and the audit trail ends up stronger than the manual process it replaced, because a machine records what it did and a busy clerk usually does not.

Credentials

Accreditations behind AI solutions

AI work reads from systems that already exist, so most of what follows is about the data platform underneath and the terms under which data moves.

Client words

What AI solutions clients say

Comments from people who run ai solutions systems day to day.

  • The handover was the part I judged them on. Configuration decisions documented with the reasoning, our administrators trained properly, and a checklist we actually worked through. We run it ourselves now, and calling them is a choice rather than a necessity.
    Head of Shared Services Multi site manufacturing group
  • What sold us was that they argued with our brief. We asked for a reporting layer and they came back saying the reporting was fine, the batch data underneath it was not, and fixing that first would cost less. That turned out to be right. Our first mock recall after go live took an afternoon instead of the better part of a week.
    Finance Director Food manufacturing group, Karachi
  • We had been through one failed implementation already, so we were sceptical of the whole category. The difference here was the migration work. Two full rehearsal loads before the real one, with a reconciliation pack we could check ourselves. Nobody had ever handed us evidence like that and asked us to sign it.
    Head of IT Wholesale distribution business
Questions

Questions about AI document processing

There is no honest menu price, because layouts, volumes and the ERP write differ completely between businesses. Rates give a reference: full stack development here is quoted at PKR 3,500 to 8,000 an hour, and the nearest published project comparator is custom web application work at PKR 200,000 to 1,000,000 and upwards. We scope and price a pilot on one document type, then quote the build.

Reported timelines give six to twelve weeks for a focused single scope go live, and a pilot on one document type fits that. Six weeks suits digital documents with agreed fields; twelve is realistic with scans, several supplier layouts or a three way match on posting. Recording the correct answers for a few hundred real documents is usually the longest item.

Accurate against what, is the question we ask back. Accuracy is measured per document type against a set of real cases with agreed correct answers, and until that set exists any figure is marketing. We publish the number we measure on your documents at the end of the pilot, per field, and we will not quote an accuracy figure before we have seen a single one of them.

Optical character recognition turns an image into text and stops there. Document processing decides what that text means: which number is the invoice total rather than the subtotal, which date is the due date, which line belongs to which purchase order. It then validates that against your master data and routes what it is unsure about to a person. The reading is the easy half.

Invoices, purchase orders, delivery notes, goods receipts, forms and structured letters, arriving as PDF, scan or email attachment. Clean digital documents are straightforward. Photographs taken on a phone, faxes, stamped and annotated copies and handwriting are all possible and all cost more, so they are assessed on your actual samples rather than assumed either way.

Wherever the design says, and it is agreed in writing before a prototype touches a real record. Which documents leave your environment, which provider processes them, in which region, what is retained and for how long, and whether the terms exclude training on your content. Logs and monitoring are part of that decision rather than an afterthought, because request content ends up in both.

Typing invoices into your ERP by hand?

Tell us what you run today and where ai document processing is causing you trouble. The first conversation is a consultation rather than a pitch.