← Back to Blog

Arabic and English AI Chatbot for Business: What Production-Grade Actually Means

The demo always works

Every vendor's demo of a bilingual chatbot goes the same way. Someone types a polite question in English, then a polite question in Modern Standard Arabic, and the bot answers both. Everyone nods.

Then the first real customer in Sharjah writes: "السلام عليكم، the 2BR في الخليج التجاري still available? بكم؟"

Three languages in one line, if you count the dialect. No punctuation the model expects. A price question with no unit named. This is what production looks like in the Gulf, and it is where most chatbots quietly fall apart.

Production-grade is not a marketing phrase. It is a specific set of properties a system either has or does not. This is the checklist we use at CloudArma before we let an agent talk to a customer.

1. It handles dialect, not just Modern Standard Arabic

Customers do not write in the Arabic of the news bulletin. They write in Gulf, Egyptian, or Levantine dialect, often with spelling that follows how words sound. A production agent understands "بكم" and "كم سعره" as the same question, and it replies in a register that feels natural rather than formal to the point of stiffness.

This is prompt work plus evaluation work. The prompt sets the register and the fallbacks. The evaluation set, built from real customer messages, proves it holds.

2. It follows the customer's language, turn by turn

Bilingual customers switch languages mid-conversation, and sometimes mid-sentence. The agent should answer each message in the language that message was mostly written in, keep names and numbers consistent across the switch, and never force a language choice up front.

Right-to-left formatting matters here too. Mixed Arabic and English with numbers and currency is a classic place for mangled output. Test it on a real phone, not just in a console.

3. Every fact comes from a tool, not from the model

The model is good at language. It is not a database. Availability, prices, opening hours, order status, and appointment slots must come from a tool call into your systems, every time. The agent should be explicitly allowed to say "let me check" and explicitly forbidden from guessing.

This one rule removes the most damaging failure mode of customer-facing AI: confident, fluent, wrong.

4. It knows what it may not say

Guardrails are the rules the agent cannot break regardless of how the conversation goes. Typical ones for a sales or support agent in the UAE:

  • No prices or discounts that are not in the system.
  • No promises about delivery, approval, or availability the tools did not confirm.
  • No personal data sent outbound to anyone but the verified customer.
  • No medical, legal, or financial advice beyond the company's approved statements.
  • Escalate immediately on complaints, threats, or anything that reads like a regulatory issue.

Guardrails live in the prompt, in the tool permissions, and in a checking layer that reviews outbound messages. Defence in depth, not a single instruction.

5. It is evaluated on real conversations, nightly

A prototype is tested once. A production system is tested continuously.

We build an evaluation suite from real transcripts, anonymised, with the expected behaviour for each. Did the agent answer in the right language? Did it call the tool before quoting a price? Did it escalate when it should? The suite runs every night and on every prompt change. A prompt that passes 312 of 312 cases gets promoted. One that regresses does not ship. This is the core of prompt engineering as an engineering discipline, and it is what most vendors mean when they say "we'll fix it if it breaks".

6. It has a fallback when the model provider has a bad day

Model providers have outages and rate limits. A production agent routes around them: a primary model, a fallback model, and a defined behaviour when both are unavailable, such as acknowledging the message and queueing for a human. Your customer should never see a raw error.

7. It hands off to a human with the whole story

The best agents know their limits. When a conversation needs a person, the person should receive the full transcript, the customer's language, the intent so far, and a suggested next step, on the channel they already use. A handoff that loses context is not a handoff. It is a restart, and customers hate repeating themselves in any language.

8. It is observable

You should be able to see every conversation, every tool call, every escalation, and every failure in one place, with latency and cost per conversation. Without this you are running blind, and you will find out about problems from customers instead of from dashboards.

9. It respects the data

Customer messages are personal data. A production deployment defines where transcripts are stored, for how long, who can read them, and how they are encrypted at rest and in transit. It uses the model provider's business terms that exclude training on your data. It can delete a customer's history on request. None of this is optional in 2026.

10. It is built to change

Your offering will change. Your prompts will need to change with it. A production system versions prompts like code, tests them like code, and rolls them out gradually. A prototype has one prompt in a text box that someone edits live.

The short version

If you are evaluating a bilingual chatbot for your business, ask the vendor to show you, live, on a phone:

  1. A Gulf dialect message with a mid-sentence switch to English.
  2. A price question where the answer must come from your data.
  3. A message that should trigger escalation.
  4. What the handoff looks like on your salesperson's screen.
  5. The evaluation results from last night.

Two of those five separate a demo from a system. All five separate a system from a liability.

Frequently asked questions

Which AI model is best for Arabic?

It changes every few months, and the honest answer is that model choice matters less than the prompt, the tools, and the evaluation around it. A production architecture lets you switch models without rebuilding, which is more valuable than picking the right one today.

Do we need separate agents for Arabic and English?

No. One agent with language-aware prompting is simpler to maintain and gives customers a consistent experience when they switch.

How do we measure quality in Arabic?

The same way as in English: an evaluation set built from your real conversations, scored on the behaviours that matter to you, run automatically.


Building or replacing a bilingual agent? See how ResAI handles real estate conversations in Arabic and English, or talk to us about your workflow.