Guide

Arabic document processing: what OCR and AI can read, and how to test it on your files

Modern OCR (optical character recognition) reads clean, printed Arabic well. The failures sit elsewhere: handwriting, stamps and signatures lying across the text, the dots and optional vowel marks that separate one Arabic word from another, right-to-left tables, and lines that mix Arabic, English and more than one set of digits. A vendor's language list will not tell you how its tool does on your forms. As read on 30 September 2026, Google's Enterprise Document OCR and Microsoft's Azure Document Intelligence (v4.0) list both printed and handwritten Arabic, open-source Tesseract publishes Arabic models, and Amazon Textract's documented languages do not include Arabic. The dependable way to choose is a test on about 200 of your own documents, scored field by field against a hand-checked answer key, with every field the system is unsure of routed to a person and a sample of the confident ones checked as well.

Arabic document processing. Extracting text and structured fields from scanned or photographed Arabic and bilingual Arabic-English documents with OCR and document AI models, with validation rules and human review for the fields the system cannot read with confidence.

Checked . Vendor language and location lists were read on Google Cloud's, Microsoft's and AWS's own documentation, the Tesseract list on its project documentation, and the Unicode Standard and its bidirectional algorithm on unicode.org. Vendors change which languages, handwriting and regions they support without notice, so check each list again before you shortlist. This page is not legal advice.

Why Arabic documents break OCR that works on English

Arabic OCR demonstrates well. Somebody feeds in a clean, typed page, the Arabic comes back letter-perfect, and the room concludes the problem is solved. Then the real intake arrives: a form filled in by hand, a scanned letter with a round stamp across the reference number, a bank statement photographed on a desk, an application with the name in Arabic and the address in English.

Some of what goes wrong is the script, and some of it is the paperwork. It helps to keep the two apart, because they have different fixes.

Joined letters with four shapes
Arabic is written right to left and is cursive even in print, and a letter that joins on both sides has four forms, for isolated, final, medial and initial positions [1]. A fold or a smudge that breaks a join changes what the model sees as a letter.
Dots and missing vowels
Several Arabic letters share one shape and are told apart only by their dots, and the short vowel marks are optional: in normal writing they are left out [1]. A dot lost on a poor scan can turn one valid word into another valid word, which a spelling check will not catch. Names suffer most, because there is no dictionary to fall back on.
Two directions on one line
A line of Arabic that carries a date, an account number or an English company name is bidirectional: digits and Latin script run left to right inside right-to-left text [2]. A tool can read every character correctly and still hand the fields back in the wrong order.
More than one set of digits
A document can carry Western digits, Arabic-Indic digits or the Eastern Arabic-Indic digits used for Persian and Urdu, and Unicode encodes each set as separate characters [1]. A check that accepts only 0 to 9 rejects a phone or identity number that was read correctly.
Tables that run from the right
A table on an Arabic form starts at the right-hand edge. An extraction that reads the cells left to right produces a table with every column in the wrong place and every value plausible.
Stamps, seals and signatures
Official documents are stamped and signed, often across the text that matters. A stamp over a reference number or a signature through a date needs a human reader whatever the model, and the design should expect it rather than treat it as an edge case.
Handwriting
Handwritten Arabic varies far more between writers than print does, and its joins, dots and letter shapes are all less regular. Two of the services below list handwritten Arabic. Neither listing is an accuracy figure.

What the main OCR services say they support in Arabic

This table sets out what four widely used options publish about Arabic on their own documentation, unranked. A listed language is where a shortlist starts.

Arabic support as each publisher documents it, read on 30 September 2026. Language lists change without notice.
DimensionPrinted ArabicHandwritten Arabic
Google Cloud Document AI, Enterprise Document OCR [3]ListedListed as supported
Azure Document Intelligence, Read and Layout models [4]ListedListed for v4.0; not in the v3.1 handwritten list
Amazon Textract [5]Not listed: AWS names English, Spanish, German, Italian, French and PortugueseNot listed
Tesseract, open source [6]An Arabic language model and an Arabic script model are publishedNot documented

Read the listings narrowly. A listed language means the publisher has trained and released a model for it. It tells you nothing about a stamped, photocopied form with a handwritten name on it, and the only way to find that out is to run your forms through the tool.

General-purpose models that read images belong in the same test, on the same files, scored against the same answer key. A model that writes fluent Arabic can also write a plausible value that is not on the page, so its output is checked against the image like everything else.

How to test it on 200 of your own documents

I would not sign a contract for Arabic document processing, or start a build, without this test. It is small enough to run before a decision and large enough to show where a tool breaks.

Choose the documents from the real intake
Take about 200 from the last year's intake, not from a folder of good examples. Weight them towards the document types with the most volume, then add the hard cases on purpose: handwritten, stamped, faxed, photographed on a phone, and mixed Arabic and English. Note which condition each document represents, because the report is broken down by it.
Build the answer key by hand
Two people type the fields that matter from each document independently, and a third settles the differences. That doubles the labelling effort and it is worth it: an answer key with errors in it makes every tool look worse, or better, than it is.
Decide what counts as right, field by field
Identity numbers, amounts, dates and phone numbers are right or wrong once digits and spacing are normalised. Names and addresses also need a softer measure, such as the character error rate, alongside exact match. Agree the rules before the first tool runs, so nobody tunes the scoring to the result.
Run every shortlisted tool unchanged
Same files, same answer key, default settings first, then any settings the vendor recommends. Record the version of each tool and API you ran, because the result belongs to that version.
Report by field and by condition
One overall accuracy figure hides what you need to know. Report each field by document type and condition, and show how the tool's own confidence scores line up with its errors. The useful question is whether the errors sit below a threshold you could send to a person.
Keep the set
The documents and their answer key become the regression test for every later change: a new model, a new vendor version, a new form.
Arabic document test: accuracy report outline
ARABIC DOCUMENT PROCESSING TEST: ACCURACY REPORT

Tool, version and API version tested:
Date run:
Documents: about 200, listed by document type and by condition
Answer key: typed by two people independently, differences settled by a third

1. Fields scored and the rule for each
   Field | Rule (exact match after normalising, or character error rate) | Why it matters

2. Results by field and document type
   Field | Document type | Documents | Correct | Errors | Example errors

3. Results by condition
   Printed | Handwritten | Stamped or signed over | Photographed | Mixed Arabic and English

4. Confidence against errors
   Threshold | Share of fields sent to review | Errors above the threshold

5. Failure notes
   What broke, with the document reference and an image crop

6. Decision
   Fields accepted automatically, fields always reviewed, and what to test next

Two hundred documents cannot settle small differences. If a condition such as handwritten names has twenty documents in the set, each error moves its score by five percentage points, so read the small groups as a direction rather than a measurement, and add documents where a decision depends on them.

Design the human review before you pick the tool

Every Arabic document system I would put into production sends some fields to a person. The design question is which fields, and how the reviewer sees them.

Route by field, not by document
A document with nine confident fields and one doubtful one should cost the reviewer one check, not ten. Route each field on the tool's confidence and on the rules below.
Validate before a person sees it
Check digits where an identifier has them, dates that exist, amounts that add up, and fields that must agree with each other, such as a Hijri and a Gregorian date for the same event. A field that fails a rule goes to review however confident the model was.
Show the image beside the value
The reviewer sees the crop of the page the field came from next to the extracted value, rendered right to left, and corrects the value rather than retyping the document.
Audit a sample of the confident fields
A review queue fed only by low confidence tells you nothing about the errors the system was sure of. Check a random sample of accepted fields as well.
Feed the corrections back
Every correction is a labelled example. Kept with its image, it grows the test set and shows whether the next version is actually better.

This is the agent-proposes, verifier-gates pattern applied to paper: the model proposes a reading, and deterministic checks and a person decide whether it is accepted.

Where the documents are processed

Scanned forms are dense with personal data: names, identity numbers, addresses, signatures. With a cloud OCR service, each image goes to wherever that service runs, which is not necessarily where your storage sits.

Check the service's own location list rather than the cloud provider's region map. As read on 30 September 2026, Google's Document AI lists its locations as the US and EU multi-regions and seven single regions with limited support, none of them in the Middle East [7], and Amazon Textract's endpoint list has no Middle East region [8]. Microsoft publishes Azure's availability by region separately, so check Document Intelligence there for the region you intend to use.

Whether documents may leave the country depends on the data, the sector and the jurisdiction, and our guide to AI data residency in the Gulf sets out which instruments decide it. Where the answer is no, the options are a service that runs in-country or a model you host yourself on in-region compute, with the model operations that come with it. This page is not legal advice.

Where 1AYM fits

1AYM has built Arabic-language AI, including bilingual Arabic-English search and OCR for Arabic documents. Where you are still choosing a tool, the AI Opportunity & Feasibility Sprint is where we would run the 200-document test: the shortlist, the answer key, the accuracy report and a build-or-buy recommendation you can put in front of whoever signs it off.

Once the decision is made, our production AI systems work builds the pipeline around the chosen tool: extraction, validation rules, the review queue, audit logging and the regression suite, running where your data is allowed to be processed. The assessment platform we run for a government-accredited EdTech in the Middle East is built the same way: cases the system is unsure of go to a person, and a random sample of the confident ones is audited too. Gulf clients can contract with 1AYM locally through its UAE entity, and if the job is already scoped, we can resource it on contract from the collective of associates who work with 1AYM, held to the same standard. If you have a document backlog and a decision to make, book a call or email us a description of the forms.

For engineers: normalisation, text order, digits and API versions to check

The details that decide whether a correct reading survives into your database, each with its source. Re-read the vendor pages before you rely on them.

Normalise, and keep the original
Unicode encodes Arabic-Indic digits at U+0660 to U+0669 and Eastern Arabic-Indic digits at U+06F0 to U+06F9 [1]. Map both to ASCII digits before validation. Remove tatweel (U+0640), the character used to stretch text for justification [1], before matching. Store the extracted string unchanged next to the normalised one.
Presentation forms from PDFs
Text pulled from born-digital PDFs sometimes arrives as Arabic presentation forms, contextually shaped glyphs that Unicode says should not be used in general interchange [1]. Compatibility normalisation (NFKC) maps them back to ordinary letters; without it, search and matching fail on text that looks correct on screen.
Logical order and display order
Bidirectional text is still interpreted in logical order, and only the display is reordered [2]. Check which order each engine returns, test lines that mix Arabic, Latin script and digits, and never fix ordering by reversing strings.
Tables from coordinates
Rebuild the rows and columns of a right-to-left table from the cell coordinates the engine returns rather than from the order of the text, and test merged cells on purpose.
Language hints on bilingual pages
Microsoft advises against passing a language code to Document Intelligence unless you are sure of the language, warning that otherwise the service may return incomplete and incorrect text [4]. On mixed Arabic and English pages, let detection run, and test both ways.
Pin the version you tested
Azure Document Intelligence lists handwritten Arabic for v4.0 and not in its v3.1 handwritten list [4], so the API version is part of the result. Record it with every run of the test set.
Measures
Exact match after normalisation for structured fields; character error rate for names, addresses and free text; the share of fields sent to review; and errors above the review threshold, which is the number that carries the risk.
Processing location as evidence
Set the processor location explicitly where the service allows it [7], and log it with each request, so the residency answer is a record rather than a setting somebody remembers.

Sources

  1. [1]The Unicode Standard, Version 18.0.0, Chapter 9, Arabic, read 30 September 2026
  2. [2]Unicode Standard Annex #9, Unicode Bidirectional Algorithm, revision 52 (1 September 2026), read 30 September 2026
  3. [3]Google Cloud, Document AI processor list, Enterprise Document OCR supported languages (last updated 24 September 2026), read 30 September 2026
  4. [4]Microsoft Learn, Azure Document Intelligence language support for the Read and Layout models (dated 18 April 2026), read 30 September 2026
  5. [5]AWS, Amazon Textract Developer Guide, Best practices (supported languages), read 30 September 2026
  6. [6]Tesseract documentation, traineddata files for version 4.00 and later (Arabic language and script models), read 30 September 2026
  7. [7]Google Cloud, Document AI regional and multi-regional support (last updated 24 September 2026), read 30 September 2026
  8. [8]AWS General Reference, Amazon Textract endpoints and quotas, read 30 September 2026

Arabic OCR questions

Can OCR read handwritten Arabic?

Some services now list it. As read on 30 September 2026, Google's Enterprise Document OCR marks handwritten Arabic as supported, and Azure Document Intelligence lists it for its v4.0 Read and Layout models but not in its v3.1 handwritten list. Handwriting is where results vary most between writers and forms, so treat those listings as a reason to test, and test on your own handwritten documents before relying on it.

Does Amazon Textract support Arabic?

Not according to AWS's own documentation as read on 30 September 2026, which names English, Spanish, German, Italian, French and Portuguese as the languages Textract supports. Check the current list before ruling it in or out, because vendors add languages without notice.

What accuracy should we expect from Arabic OCR?

Nobody can give you an honest figure without your documents. Accuracy on clean printed pages says little about stamped, handwritten or photographed forms. Build a test set of about 200 of your own documents with a hand-checked answer key, and measure each field by document type and condition.

Can a large language model replace OCR for Arabic forms?

Put it in the same test rather than assume either way. A model that accepts images can be asked to read a page, but a model that writes fluent text can also produce a plausible value that is not on the page. Whatever reads the document, validate the fields and send the doubtful ones to a person.

How should we handle Arabic-Indic digits and vowel marks?

Normalise for matching and keep the original for the record. Map Arabic-Indic and Eastern Arabic-Indic digits to Western digits before validating a number, strip the optional vowel marks and the tatweel stretching character for search, and store the text exactly as extracted alongside the normalised copy.

More in this topic

  • What Abu Dhabi's health rules ask of an AI system that reads patient records and claims.

  • What Saudi Arabia's personal data law asks when forms full of personal data pass through an AI system.

Further

  • AI strategy and roadmap · The sprint where a document test and the build-or-buy decision sit, before any platform is bought.
  • Data strategy for AI · What to fix in documents and databases before an AI system reads them.

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM