Guide

Bilingual Arabic-English enterprise search with RAG: what breaks and how to test

Bilingual Arabic-English enterprise search lets staff ask a question in either language and get an answer drawn from documents in both, with the source cited. It is usually built as retrieval-augmented generation (RAG): a search step finds the passages, and a language model writes the answer from them. The search step is where it fails. Arabic attaches prefixes and suffixes to words, spells the same word several ways and usually leaves the vowels out, so keyword search tuned for English misses documents that are plainly relevant. Names have one spelling in Arabic and several in English, and the answer to an English question may sit only in an Arabic circular. A system that works needs Arabic-aware keyword search, a multilingual embedding model, permissions enforced at the search step, and a test set of real questions in both directions, scored by people who read both languages. The model that writes the answer is the easier part.

Bilingual Arabic-English enterprise search. An AI search and question-answering system over documents written in Arabic, English or both, which finds the relevant passages whichever language the question is asked in and answers with a citation to the source.

Checked . The script facts were read in the Unicode Standard on unicode.org, the analyser behaviour in the Apache Lucene, Elastic and Microsoft Learn documentation, and the retrieval security guidance on OWASP's site. Search products change their analysers and defaults between versions, so check the documentation for the version you run. This page is not legal advice.

Why search tuned for English misses Arabic documents

The demo usually goes well. Someone asks the assistant a question in English, the answer comes from an English policy, and it is right. The trouble starts when a colleague asks the same question in Arabic, or when the only document that answers it is an Arabic circular from three years ago, or when the question mixes both languages because the system it asks about has an English name.

Each of those failures has a mechanical cause, and nearly all of them sit in the search step, before the language model has seen a word. A model cannot answer from a passage it was never given.

Words carry their prefixes and suffixes
Arabic attaches the article, some prepositions, conjunctions and pronouns directly to the word, so one written word can carry what English spreads over three or four. An index that splits text on spaces stores each combination as a separate term, and a search for the bare word misses them. Search engines handle this with light stemming, which strips the common prefixes and suffixes, and Lucene's Arabic analyser does exactly that [2].
One word, several spellings
The same word turns up with and without the hamza on its alef, ending in teh marbuta or in heh, with a dotless yeh or a dotted one, and with or without the short vowel marks, which normal writing leaves out [1]. Lucene's Arabic normaliser folds those variants together and removes the tatweel stretching character as well [3]. Without that step, two spellings of one word are two different words to the index.
Names in two scripts
A person, a company or a place has one spelling in Arabic and several in English, and your documents will use all of them. Neither stemming nor normalisation connects them. That takes a list of the names that matter, in both scripts, with somebody responsible for keeping it current.
The answer is in the other language
An English question whose answer exists only in an Arabic document shares no words with it, so keyword search finds nothing. Crossing languages is the job of a multilingual embedding model, which places a passage and its translation close together, or of translating the question before searching.
Questions that switch language
People write Arabic questions with English system names, product codes and acronyms in them. Each part needs the analysis for its own language, and a query treated as all Arabic or all English loses one of them.
Digits and extracted text
Numbers appear in Western digits and in Arabic-Indic digits, which Unicode encodes as separate characters [1], so a reference number typed in one form misses the other unless both are mapped to the same digits first [7]. Text pulled from PDFs can also arrive as Arabic presentation forms, which Unicode says should not be used in general interchange [1] and which do not match the ordinary letters a user types.

What a bilingual system needs, layer by layer

Nothing here is exotic. These are decisions that cost little at the start and a great deal to change once the archive has been indexed the wrong way, which is why I would settle them before the first document goes in.

Two ways to search, combined
Run keyword search and vector search side by side. Keyword search, with an Arabic analyser on Arabic text and an English one on English, finds exact terms, reference numbers and names. Vector search with a multilingual embedding model finds passages that mean the same thing in either language. Merge the two result lists and put a multilingual reranker over the top. Either route alone misses what the other finds.
Chunks that follow the document
Split documents at their own headings, clauses and articles, rather than at a fixed character count that cuts an Arabic sentence in half. Where a page carries both languages side by side, make an Arabic chunk and an English chunk of the same clause and link them, rather than one chunk that mixes the two.
A translation is one document
Policies, circulars and contracts in the Gulf often exist in both languages. Index the pair as one document with two language versions, so the assistant does not cite the same clause twice as two sources, and record which version your organisation treats as authoritative, because the two sometimes differ.
Metadata you will need later
Language, version, issuing body, effective date and classification on every chunk. An answer drawn from a superseded circular is wrong however well it is written, and the effective date is what lets the system prefer the current one.
Permissions at the search step
The search returns only what the person asking is already allowed to open, using the access lists of the systems the documents came from. OWASP's guidance on retrieval systems puts fine-grained, permission-aware vector stores and strict partitioning of the data first among its mitigations [5]. An instruction in the prompt telling the model to ignore documents the user cannot see is not an access control.
The language of the answer
Answer in the language of the question, cite the passage in the language it was written in, and say so when the source is in the other language, so a reader who needs the authoritative wording knows where it is.

How to test it before anyone relies on it

I would not put a bilingual assistant in front of staff, let alone customers or the public, without this test. It needs time from people who know the documents and read both languages, and it is the only evidence that the system works in both languages rather than mostly in one.

Collect real questions
Take about 200 from search logs, helpdesk tickets and the people who answer questions today, in the language they were actually asked in. Fill every cell of the table, then weight the rest of the set towards the questions people ask most.
Write the answer key by hand
For each question, the documents and passages that answer it and a short reference answer. People who read both languages write it, and a second person checks it.
Score the search and the answer separately
First, whether the right passage was among those the model saw. Then whether the answer is correct, whether the passage it cites supports it, and whether it declined when there was nothing to find. A wrong answer with the right passage in hand is a generation problem; a wrong answer with the passage missing is a search problem, and the two have different fixes.
Have the Arabic answers judged in Arabic
A native Arabic reader scores the Arabic answers for correctness and for whether they read the way your organisation writes. An automated judge can help screen the results, but it should not be the only reader of either language.
Report by cell
One overall accuracy figure hides the cell that fails. Report each cell on its own, and treat a single permission leak as a failure of the whole test.
Keep it as a regression suite
Rerun the set on every change to the analysers, the embedding model, the chunking or the model that writes the answers. A change that improves English and quietly makes Arabic worse shows up here and nowhere else.
The cells of a bilingual test set. Fill every one and score each on its own.
DimensionWhat a question looks likeWhat it catches
English question, English answerThe case most demos showWhether the basics work at all
Arabic question, Arabic answerA question typed the way staff write, usually without vowel marksMissing normalisation and stemming
English question, answer only in ArabicA policy that exists only as an Arabic circularWhether the search crosses languages
Arabic question, answer only in EnglishA technical standard or a contract annex written in EnglishThe same, in the other direction
Mixed-language questionAn Arabic question naming an English system or product codeQueries analysed as one language
Names and numbersA person, supplier or reference number in its other spellings and digit formsTransliteration gaps and unmapped digits
No answer existsA plausible question your documents do not answerAn assistant that invents an answer rather than saying it has none
AccessA question from a user who may not open the document that answers itPermission leaks, where the only acceptable count is zero
Bilingual search test: plan and report outline
BILINGUAL ARABIC-ENGLISH SEARCH TEST: PLAN AND REPORT

System tested: search configuration, embedding model, reranker and answer model, with versions
Date run:
Documents in scope: sources, languages, count, date range
Questions: about 200, each tagged with its cell and the language it was asked in
Answer key: written by readers of both languages, checked by a second person

1. Cells and question counts
   EN to EN | AR to AR | EN question, AR answer | AR question, EN answer
   Mixed-language | Names and numbers | No answer exists | Access

2. Search results by cell
   Cell | Questions | Right passage retrieved | Missed | Example misses

3. Answer results by cell
   Cell | Correct | Citation supports the answer | Declined correctly | Invented

4. Access tests
   User role | Questions | Documents returned that the user may not open (must be zero)

5. Arabic answers, judged by a native reader
   Correct | Reads as your organisation writes | Notes

6. Failure notes
   What broke, with the question, the expected passage and what came back

7. Decision
   What goes live, for whom, in which language, and what is tested next

Two hundred questions spread over eight cells gives each cell about twenty-five, so a single question moves a cell's score by four percentage points. Read the small cells as a direction rather than a measurement, and add questions where a decision rests on them.

Where the index, the embeddings and the model run

A search system makes copies. The index holds the text of every document it covers, the vector store holds embeddings computed from that text, and every question sends the retrieved passages to whichever model writes the answer. OWASP warns that inadequate access controls on embeddings can expose the sensitive information in them [5], so each copy needs its own answer on location and access, not only the document store.

Whether any of those copies may leave the country depends on the data, the sector and the jurisdiction. Our guide to where Gulf law requires AI data to stay sets out which instruments decide it, and which cloud regions in the Gulf actually run models. This page is not legal advice.

Where 1AYM fits

1AYM has built Arabic-language AI, including bilingual Arabic-English search and OCR for Arabic documents. If you are deciding whether to put an assistant over a two-language archive, or working out why the one you have answers well in English and badly in Arabic, the AI Opportunity & Feasibility Sprint is where we would scope that test against your documents and give you a build-or-fix recommendation.

Once the decision is made, our production AI systems build covers the rest: the analysers and embeddings, retrieval that respects permissions, the answer and citation layer, audit logging, and the full test set run as the regression suite, all where your data is allowed to be. For a government-accredited EdTech in the Middle East we run a production estate whose database was replatformed into Google Cloud's Doha region to meet Gulf data-residency requirements, with more than 1,400 automated tests across the platform. Gulf clients can contract with 1AYM locally through its UAE entity, and if the job is already scoped, we can resource it on contract from the collective of associates who work with 1AYM, held to the same standard. If you have documents in two languages and a question about searching them, book a call or email us.

For engineers: analysers, normalisation, digits, hybrid retrieval and permission filters

The settings that decide whether an Arabic passage is found at all, each with its source. Re-read the documentation for the version you run before relying on it.

The Lucene Arabic chain
Lucene's ArabicAnalyzer runs StandardTokenizer, then lower-casing, decimal digit folding, stop words, Arabic normalisation, an optional keyword marker and Arabic light stemming, the last following the paper Lucene cites, Light Stemming for Arabic Information Retrieval [2]. Elasticsearch documents the same chain for its built-in arabic analyser and shows how to rebuild it with your own stop and keyword lists [4].
What the normaliser folds
Hamza forms of alef to bare alef, teh marbuta to heh, alef maksura to yeh, and it strips the harakat and tatweel [3]. Apply the same analyser at index and query time, and keep the original text in a stored field for display and citation.
Two analysers to compare
Azure AI Search offers ar.microsoft and ar.lucene for Arabic fields, and Microsoft advises comparing the two with its Analyze API to see the tokens each produces [6]. Whichever engine you use, run your own names, codes and mixed-language strings through the analyser before indexing anything.
Digits
Arabic-Indic digits sit at U+0660 to U+0669 and Eastern Arabic-Indic at U+06F0 to U+06F9 [1]. Elastic's decimal_digit filter converts every Unicode decimal digit to 0 to 9 [7]. Index reference numbers, identity numbers and amounts in a keyword field after that folding, not only in the analysed text.
Presentation forms
Arabic presentation forms from PDF extraction should not be used in general interchange [1]. Apply compatibility normalisation (NFKC) before analysis, so contextually shaped glyphs become the ordinary letters a user types.
Hybrid retrieval and reranking
Keep language-specific keyword fields beside a multilingual vector field, fuse the two result lists (reciprocal rank fusion is a simple place to start), then rerank the fused top results with a cross-encoder that has seen Arabic. Check the embedding model with known translation pairs from your own documents: the Arabic and English versions of one clause should retrieve each other.
Permission filters before ranking
Carry the source system's access lists into the index and filter inside the query, before ranking, rather than dropping results after the top k are chosen, which returns short and uneven result sets. OWASP's first mitigation for vector stores is fine-grained, permission-aware access and strict logical partitioning [5]. Test it with users who hold different roles.
Measures
Recall at k for the passages in the answer key, by cell; answer correctness and citation support, by cell; the share of unanswerable questions declined; and permission leaks, where the target is zero. Log the retrieved passage IDs with every answer, so an answer can be traced to what the model was shown.

Sources

  1. [1]The Unicode Standard, Version 18.0.0, Chapter 9, Arabic, read 30 September 2026
  2. [2]Apache Lucene 10.5.1 API documentation, ArabicAnalyzer, read 30 September 2026
  3. [3]Apache Lucene 9.12.0 API documentation, ArabicNormalizer, read 30 September 2026
  4. [4]Elastic documentation, language analyzers, Arabic analyzer, read 30 September 2026
  5. [5]OWASP Top 10 for LLM Applications 2025, LLM08:2025 Vector and Embedding Weaknesses, read 30 September 2026
  6. [6]Microsoft Learn, Add language analyzers to string fields, Azure AI Search (dated 16 June 2025), read 30 September 2026
  7. [7]Elastic documentation, decimal digit token filter, read 30 September 2026

Bilingual search questions

Can one AI assistant answer in both Arabic and English?

Yes, if the search underneath it finds passages in both languages. The model that writes the answer is rarely the weak point. Test the whole system on real questions in each direction, including English questions whose only answer is in Arabic, before deciding it works.

Should we translate our Arabic documents into English and search only in English?

It is a reasonable first step for a small archive, but it adds a translation error to every search, loses the authoritative wording, and costs more with every new document. I would search both languages directly and translate only the question, if at all.

What is RAG, in plain terms?

Retrieval-augmented generation: the system first searches your documents for the passages relevant to a question, then gives those passages to a language model to write the answer, with citations. The model answers from your documents rather than from what it learned in training.

How many test questions do we need?

About 200 real questions is enough to see where a system breaks, provided they cover every cell: each direction between the languages, mixed-language questions, names and numbers, questions with no answer, and users with different access. Add questions where a decision rests on a small cell.

Do dialects and Arabic typed in Latin letters matter?

They matter for chat, where people type the way they speak. Formal documents are usually written in Modern Standard Arabic, so the larger gap is usually between how people ask and how documents are written. Put questions typed the way your users type into the test set and measure it.

More in this topic

  • Scanned Arabic forms have to be read before they can be searched: where OCR fails and how to test it.

  • What Saudi Arabia's personal data law asks of a system that indexes documents full of personal data.

Further

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM