OCR Development Services

Custom OCR for the documents that keep breaking generic extraction.

We benchmark your real scans and photos, then build the smallest reliable path from image to traceable text or fields. Capture checks, preprocessing, provider selection, field-level evaluation, review, and one integration are designed as one system.

Bring representative documents, the values you need, and the system receiving them. Leave with a buy, configure, build, narrow, or wait recommendation.

Recorded OCR-assisted workflow

Musgrave Group logo

Musgrave Group

RaftLabs delivered a receipt-entry and campaign-decision platform for SuperValu and Centra in Northern Ireland. Customers submitted real receipt photos rather than a controlled batch of clean scans.

1,610
receipt submissions at the reporting snapshot
1,062
users recorded in the campaign report

The brief

Start with what is not working.

Good software decisions begin with the constraint, not a list of features or a preferred technology.

01

The demo reads a clean PDF, but glare, folds, stamps, tables, handwriting, or a supplier's new layout changes the result.

02

A confidence score looks reassuring, yet nobody has tested whether the correct value was attached to the correct field.

03

People still compare every extraction with the source because the system cannot explain, reject, or route an uncertain result.

Plain answer

OCR development turns text and named fields in images, scans, and PDFs into machine-readable output. Custom OCR is useful when a standard product or API cannot reliably handle the real document quality, layouts, fields, languages, or operating controls your workflow requires.

What to remember

  • Benchmark maintained OCR products and APIs before funding custom model work.
  • Measure important fields by document cohort; a single character or document accuracy score can hide the error that matters.
  • Treat confidence as a routing signal, not proof that an extracted value is correct.
  • Start with one document cohort, one output contract, and one destination before adding more formats.

Fit

Custom OCR should solve a recognition constraint a maintained product cannot.

The business case is strongest when the same document cohorts recur, the required output is precise, and an existing provider has been tested against the difficult examples rather than dismissed on assumption.

A fit
01

Scans, photos, forms, receipts, labels, or legacy documents contain valuable text or fields that an existing product misses repeatedly.

02

The team can provide representative documents, target output, known failures, downstream rules, and people who can establish ground truth.

03

Capture quality, layout, language, vocabulary, source coordinates, privacy, latency, or deployment creates a defensible need for custom work.

Not a fit
01

The documents are already machine-readable or a maintained connector produces acceptable output with less ownership burden.

02

The real need is approval, routing, reconciliation, or accounts-payable automation rather than recognition itself.

03

Nobody can define the important fields, acceptable error, representative document mix, or person responsible for exceptions.

Custom does not have to mean training a model. It may mean better capture, preprocessing, provider orchestration, field mapping, validation, evidence, and integration around a maintained recogniser.

The text was readable. The wrong number still reached the record.

A receipt contains a transaction number, loyalty identifier, subtotal, tax, and final total. The recogniser reads every character. The application then selects the number nearest a familiar label, but glare has shifted the detected region and the label belongs to the line above. The output is fluent, properly formatted, and wrong.

That failure is easy to miss when success is reported as a page-level accuracy percentage. The buyer does not need every character equally. They need the correct merchant, date, identifier, amount, or line item attached to the correct field, in the format the receiving system accepts. A single false acceptance can matter more than a hundred harmless punctuation errors.

Production OCR therefore starts before recognition and ends after it. Capture quality, page orientation, layout, vocabulary, normalisation, validation, source coordinates, thresholds, and review determine whether a model output becomes a useful product capability or another queue people quietly recheck.

Is custom OCR the right fix, or is a product enough?

Start with the smallest maintained option that can meet the requirement. Owning a custom recognition path creates evaluation, monitoring, provider, and support work. That cost is justified only when the documents or product experience create a meaningful constraint.

Choose by the missing capability

ScenarioRecommendationReason
You need searchable text from common, good-quality PDFs or images.Use an OCR product or APIA maintained recogniser is usually faster and cheaper than custom development when plain text is the useful output.
A provider recognises the document, but capture, mapping, normalisation, or delivery is incomplete.Configure and integrateKeep the provider and build only the missing product controls, output contract, or connection.
Real inputs repeatedly defeat existing tools on valuable fields, layouts, scripts, or operating constraints.Build a custom OCR pipelineBenchmark recognisers, then add the smallest combination of preprocessing, layout logic, models, rules, evidence, and review that closes the measured gap.
The outcome is a validated record, approval, case, or posted transaction.Design an IDP workflowRecognition is only one component; classification, business validation, exception handling, write-back, and reconciliation define the service boundary.

Use intelligent document processing when the full document-to-record path is unresolved. For supplier invoices, purchase-order checks, tax, approvals, posting, and payment controls, the more precise service is invoice processing automation.

Production scope

What belongs in a custom OCR release?

The recogniser is one replaceable part. The product has to preserve what arrived, produce a defined output, show where each important value came from, and fail in a way people can operate.

  • 01
    Capture and input quality
    Guide the user or device before submission. Detect blur, glare, cutoff, skew, rotation, low resolution, compression, duplicate pages, and missing regions where possible. Rejecting an unreadable image early is often safer than asking a stronger model to guess.
  • 02
    Preprocessing and layout
    Deskew, crop, denoise, normalise contrast, separate pages, preserve reading order, locate tables or regions, and route known layout families where the benchmark shows a benefit. Preprocessing should be measured, because an aggressive image transform can also remove faint but important content.
  • 03
    Recognition and provider strategy
    Compare maintained OCR services, document models, vision-language models, and open components against the actual cohort. Use a specialised or trained model only when the field-level improvement, privacy boundary, latency, or unit economics justify the maintenance it adds.
  • 04
    Structured output and source evidence
    Return the text, fields, tables, types, units, dates, coordinates, page references, normalised values, confidence, and pipeline version the caller needs. Keep original and normalised values separate so a reviewer can see whether the error came from recognition or transformation.
  • 05
    Validation and human review
    Check formats, arithmetic, known identifiers, cross-field relationships, and reference data around the recogniser. Route uncertain or consequential fields with the source region, proposed value, failed rule, and correction action together instead of sending an unexplained score to a shared inbox.
  • 06
    Integration and observability
    Connect one capture route and destination. Preserve document identity, protect files, handle retries and duplicates, and monitor quality, review rate, failures, latency, provider cost, and changes in layouts or image sources. Successful extraction is incomplete until the caller receives and can reconcile the result.

Recorded proof

A smaller identifier space made one receipt decision easier to constrain.

For a Musgrave receipt-scanning campaign, the useful result was not a text dump. The platform had to connect a customer-submitted receipt to the correct store and campaign rules, stop an exact recorded transaction from being entered twice, and keep rejected or unresolved decisions visible to operators.

During development, long internal store identifiers made the model's choice harder to constrain. The team replaced them with short sequential integers in the model request, then mapped the selected value back to the application's internal identifier. The team associated that change with store matching moving from roughly 80% to near 99%.

That figure has a strict limit. The retained record does not include the sample, scoring method, error breakdown, or independent audit. It is not character accuracy, field accuracy across the receipt, or a benchmark another project can inherit. The lesson is narrower and more useful: reference-data design and application constraints can improve a document decision without pretending the answer is always a new model.

How should OCR quality be judged before release?

An evaluation should resemble the documents and decisions the system will meet after launch. The result is a boundary: what can pass, what must stop, and what has not yet been proven.

  • 01
    Create ground truth for the output the business uses

    Label the required text, fields, tables, relationships, and acceptable normalisation on representative documents. Keep a held-back set so the examples used to tune prompts, rules, or models are not the same examples used to approve the release.

  • 02
    Report important fields and cohorts separately

    Split results by source, layout, capture quality, language, and consequential field. Google Cloud's Document AI evaluation guidance reports precision, recall, and F1 against annotated test documents and notes why one generic accuracy measure can be misleading when labels are optional or repeated.

  • 03
    Measure false acceptance as well as missed extraction

    A blank or rejected field is visible. A plausible wrong value that passes automatically is quieter and may be more expensive. Set thresholds by consequence, then record correct automatic acceptance, incorrect acceptance, rejection, and human correction rather than celebrating coverage alone.

  • 04
    Test whether confidence predicts real error on this cohort

    Provider confidence can help route results, but it needs calibration against labelled examples. Thresholds trade precision against recall and should not substitute for field rules, source evidence, or human review in a consequential path.

  • 05
    Track operating work, not only model quality

    Measure recapture, processing failures, review rate, correction time, latency, unit cost, destination rejection, and new-layout drift. A recogniser can score well while the complete product still creates more checking than it removes.

How it works

Close one recognition risk before opening the next.

Each phase produces a decision a buyer can inspect. More documents, languages, fields, or providers enter only after the current boundary is measured.

  1. 01
    Frame

    Define the document and output contract

    Which recognition errors actually change a business result?

    Trace the document from capture to the caller that consumes the output. Separate fields that must match exactly from values that permit a tolerance, and distinguish recognition from mapping, normalisation, and business validation.

    Decision produced

    A document-cohort map, capture conditions, required output, source evidence, normalisation rules, costly errors, and pass, review, or reject conditions.

    Risk closed

    Optimising page-level text while the identifier, amount, date, table, or source region that matters remains undefined.
  2. 02
    Compare

    Benchmark existing recognisers on real inputs

    Does a maintained product already solve enough of the problem?

    Include common documents, awkward images, known failures, rare but costly fields, and held-back samples. Record where quality comes from and whether the improvement survives outside the examples used during development.

    Decision produced

    A field-level comparison of viable products, APIs, models, preprocessing, latency, cost, deployment, and expected review load, with a buy, configure, build, narrow, or stop recommendation.

    Risk closed

    Funding custom model work before testing a simpler maintained route, or judging every option on clean documents from one layout.
  3. 03
    Control

    Build only the missing recognition path

    Which additions close the measured gap without creating unnecessary ownership?

    Prefer replaceable components and explicit contracts. Keep the original document and source coordinates available, put deterministic checks around probabilistic output, and make failure reasons visible to the people who will operate the release.

    Decision produced

    A bounded capture, preprocessing, recognition, output, evidence, review, and integration path with versioned acceptance tests.

    Risk closed

    Combining several opaque models and rules until the demonstration works, then leaving nobody able to explain or change the pipeline.
  4. 04
    Operate

    Release one cohort and watch the exceptions

    Does the live workflow remove more checking than it creates?

    Start with one controlled source or document family. Add confirmed failures to a versioned regression set, rerun it when the model or provider changes, and expand only after the first cohort meets the agreed product and operating targets.

    Decision produced

    A monitored release with field quality, false acceptance, review effort, recapture, failure, latency, cost, drift, and rollback evidence.

    Risk closed

    Expanding to more sources and layouts while corrections, new document variants, and provider changes remain invisible.

Control

What should remain visible after the OCR feature ships

These conditions keep a probabilistic recognition component testable, replaceable, and supportable when documents, providers, and business rules change.

  • 01
    Every important field can point back to its source

    Retain the original value, normalised value, page or region, validation result, recogniser version, and final action as the workflow requires. The caller should be able to explain whether a problem came from the image, recognition, field mapping, normalisation, or destination.

  • 02
    The evaluation set is versioned without becoming an uncontrolled archive

    Record cohort, expected output, consent or usage basis, retention, and access. Add confirmed production failures carefully, redact where appropriate, and keep sensitive documents within the agreed storage and provider boundary.

  • 03
    A provider change reruns the same consequential tests

    Pin or record provider, model, prompt, parser, preprocessing, schema, and rules. Compare the next version on the held-back and regression sets before moving traffic, with a rollback or manual route available if a critical cohort regresses.

  • 04
    The integration can retry without duplicating the result

    Use stable document and request identifiers, idempotent operations where supported, visible failure states, and a repair path. Monitor whether the caller accepted the result, not merely whether the OCR endpoint returned a successful response.

Every OCR project starts at $9,500.

The first paid phase is deliberately bounded. It should prove whether custom OCR is justified and close the most expensive uncertainty before a production build is priced.

The 30-minute call comes first and costs nothing. If a maintained product already works, the sample is not representative, or the wider problem belongs to IDP, the recommendation can stop or change direction there.

What the first phase can be

  1. 01

    Document cohort and error map

    Inventory the real sources, layouts, capture conditions, languages, required output, costly errors, and gaps in the available samples.

  2. 02

    Provider and model benchmark

    Compare viable OCR products, APIs, models, and preprocessing on held-back documents, with important fields and cohorts reported separately.

  3. 03

    Capture, review, and output design

    Define what the product rejects early, what passes, what stops for review, what evidence is kept, and what one destination receives.

  4. 04

    Buy, configure, build, narrow, or stop decision

    Receive the evidence, recommended boundary, unresolved risks, production estimate, and conditions that justify another phase.

This is not a promise to deliver a complete production OCR platform for $9,500. It is the smallest useful phase for replacing a polished demonstration with evidence a buyer can act on.

Starting investment

$9,500

Minimum project scope. The sample boundary, deliverable, scoring method, acceptance criteria, ownership, price, and exclusions are written down before the phase starts.

Price held for the phase

The agreed phase price does not move unless you approve a material change in document cohort, output, destination, or acceptance criteria.

Client-controlled evidence

Project-specific documents, labels, expected values, code, configuration, and client accounts remain under client control, subject to external licence and provider terms.

No accuracy promise from demo data

The phase reports what was tested, how it was scored, where it failed, and what remains unknown before a production result is claimed.

OCR development questions buyers ask

OCR development services design, build, and integrate software that converts text or named fields in images, scans, and PDFs into machine-readable output. The work can include capture-quality checks, image preprocessing, OCR or vision-model selection, layout handling, structured field mapping, validation, source coordinates, human review, evaluation, monitoring, and delivery to an application or business system.

OCR recognises characters, text, layout, or fields from a document image. Intelligent document processing owns the wider document-to-record workflow: intake, classification, extraction, validation, exception handling, approval, write-back, and reconciliation. Choose OCR when recognition is the missing component inside a workflow you already own; choose IDP when the whole operating path needs to be designed.

Start with a maintained product or API when it supports your document types, languages, deployment boundary, output, and economics. Custom OCR development is justified when capture conditions need product-level controls, existing tools repeatedly miss valuable fields, layout or vocabulary needs specialised handling, several recognisers must be combined, or evidence and integration requirements do not fit an off-the-shelf route. The first phase should benchmark these choices rather than assume custom training.

No responsible supplier should promise perfect results across unseen documents. Quality changes with image resolution, blur, glare, skew, handwriting, language, layout, typography, tables, and the definition of a correct match. A production system should state what was tested, report the important fields separately, reject or review uncertain results, and keep measuring new document cohorts after release.

Create ground truth from representative documents, keep a held-back test set, and score the output required by the business. Depending on the job, that can include character or word error rate, exact field match, normalised field match, precision, recall, false acceptance, false rejection, coverage, and review rate. Report results by important field and document cohort so a strong average cannot hide a weak total, identifier, or date.

Yes. Confidence describes a model's estimate, not an independent verification of the business value. A recogniser can be confident about the characters it saw while the application maps them to the wrong field, normalises them incorrectly, or selects a plausible value from the wrong region. Thresholds need to be tested against labelled documents and combined with validation, source evidence, and review where the consequence warrants it.

Sometimes, but those are separate evaluation problems rather than checkboxes. Handwriting varies by writer and context; tables require row, column, and merged-cell structure; multilingual documents may need script detection, language hints, and domain vocabulary. We test the actual cohorts and required output before committing to an automation rate or model approach.

There is no useful universal number. The set must cover the real variations: sources, layouts, devices, image quality, languages, tables, handwritten regions, uncommon fields, and known failures. We inventory those cohorts, identify gaps, and reserve documents that are not used for tuning so the evaluation does not reward memorisation of the development set.

Yes, when the destination has an API, event interface, import route, or controlled database boundary. The output contract should define field types, coordinates or source references, confidence, normalisation, document identity, duplicate behaviour, error states, retries, permissions, and what happens when the destination refuses or partly accepts a result.

The scope names where source files, extracted values, model requests, corrections, and logs are stored; which providers process them; who can access them; and when each artefact is deleted. Encryption, role-based access, redaction, regional processing, private deployment, and client-controlled provider accounts can be included where the use case requires them. Legal or regulatory compliance is only claimed after the relevant controls are verified.

Every RaftLabs project starts at $9,500. For OCR, the first phase may cover a document-cohort inventory, field and error map, representative evaluation set, provider benchmark, capture or review design, and a buy, configure, build, narrow, or stop recommendation. A production implementation is priced separately after the evidence exposes the real boundary.

A benchmark or evidence phase can take a few weeks. A production release takes longer and depends on document variation, capture channels, field count, tables, handwriting, languages, validation, review, security, throughput, and integration. The schedule is set after representative samples and acceptance criteria replace assumptions with a measurable scope.

Work with us

Bring the documents that fail, not only the examples that work.

In a 30-minute call, we will identify the document cohort, the output that matters, and whether the sensible next move is to buy, configure, benchmark, build, narrow, or wait.

  • The first benchmark uses representative and held-back documents, with important fields reported separately.
  • The scope states what passes automatically, what stops for review, and what remains unknown.
  • Source files, model providers, retention, permissions, and destination-system boundaries are written down.
  • Project-specific code, labels, evaluation sets, client accounts, and operating notes remain under client control.