The Arabic Language Models and What You Can Actually Do With Them
Five state-backed families and a set of adaptations carry the Arabic-AI conversation, and three facts decide between them: the parameter count, what the licence actually permits, and whether the weights can be downloaded at all. This is that roster, checked against the model cards and licence texts, then the harder question of what an Arabic score measures.
Most of the field is five state-backed families and a set of adaptations
Jais came out of G42's Inception with MBZUAI and Cerebras; Jais-13b was announced on 30 August 2023. The September 2024 family release put out, the card claims, 20 models across eight sizes, 590M to 70B, though the organisation also carries 256M checkpoints the card does not count; the card separates two lineages: jais-family checkpoints pre-trained from scratch, and jais-adapted-7b, -13b and -70b adapted from Llama-2 with 32,000 Arabic tokens added to the tokeniser. The adapted 70B's config still reports LlamaForCausalLM.
Jais 2 was announced on 9 December 2025 in 8B and 70B, both from scratch: the architecture is Jais2ForCausalLM and the vocabulary 150,272 tokens, so there is no Llama or Gemma underneath. Its technical report, arXiv:2608.13580, is stamped 7 July 2026 on arXiv despite an identifier in the August block, and the abstract gives a relative training budget rather than a token count; the 2.6-trillion figure in circulation comes from Cerebras' blog.
ALLaM comes from the National Center for AI at Saudi Arabia's SDAIA. Its paper, arXiv:2407.15390, is explicit about what it covers: "four models at three different scales: 7B, 13B, and 70B models initialized by Llama-2 weights and a 7B model from scratch/random initialization". The 34B that later reached the public is not in it. Fanar is QCRI's, at HBKU in Qatar (arXiv:2501.13944): Fanar Star, 7B trained from scratch on nearly a trillion Arabic, English and code tokens, and Fanar Prime, 9B, continually trained on the Gemma-2 9B base over that same token set. Fanar 2, March 2026, moves to gemma-3-27b-pt.
TII announced Falcon Arabic on 21 May 2025, built on Falcon 3-7B, and Falcon-H1 Arabic on 5 January 2026 in 3B, 7B and 34B. Egypt's Applied Innovation Center published Karnak-6B-v1.0 and Karnak-40B-v1.0, depth-extended from Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 respectively. Around them sit AceGPT (arXiv:2309.12053), SILMA-9B on Gemma-2, and Atlas-Chat for Moroccan Darija (arXiv:2409.17912).
- Jais 2 technical report (arXiv:2608.13580)
- Cerebras, 'Jais 2: A Blueprint for Sovereign AI', 9 December 2025
- ALLaM: Large Language Models for Arabic and English (arXiv:2407.15390)
- Fanar: An Arabic-Centric Multimodal Generative AI Platform (arXiv:2501.13944)
- TII announcement of Falcon Arabic, 21 May 2025
- Karnak-40B-v1.0 model metadata (Hugging Face API)
Announced, released and downloadable are three different states
As of this writing there are no public weights for Falcon Arabic or Falcon-H1 Arabic on TII's Hugging Face organisation. The 137 repositories under tiiuae contain nothing matching "arab", and the plausible names return HTTP 401 where tiiuae/Falcon3-7B-Instruct returns 200. A 401 does not distinguish private from deleted from manually gated from never-created — Hugging Face returns it for any name it will not confirm — so this is not a claim of withdrawal, only that a reader cannot download them. TII's own 5 January 2026 announcement points readers to a chat interface at chat.falconllm.tii.ae rather than to a weights repository.
Fanar Star does not appear among QCRI's public repositories. Six Fanar checkpoints do download, all ungated and all tagged apache-2.0: the Fanar-1-9B base, its Instruct variant — 8.78 billion parameters, smaller than the 9.24 billion of the base it was tuned from, despite both carrying "9B" in the name — Fanar-2-27B-Instruct, and the Fanar-2 Diwan and Oryx models. Of the four ALLaM models the paper describes, one checkpoint is downloadable: ALLaM-7B-Instruct-preview. ALLaM 34B, which the paper does not cover, sits behind HUMAIN Chat, described in an independent evaluation of it as "a closed conversational web service".
Closure limits what can be measured. The independent evaluation of ALLaM 34B published as arXiv:2508.17378 was driven through that chat web interface: 23 prompts, five runs each, 115 outputs scored by three frontier models as judges. Its lowest category was dialect fidelity, 4.21 out of 5, against 4.74 for Modern Standard Arabic. Model or service wrapper, the two cannot be told apart from outside.
"Open" is doing four jobs here. Ungated and directly downloadable: ALLaM 7B, the Fanar checkpoints, SILMA, Atlas-Chat, both Karnak models, and a minority of the Jais repositories, jais-family-256m, -590m, -13b-chat and -30b-8k among them. Behind an automatic click-through: Jais 2 in both sizes, every full-precision jais-adapted checkpoint, most of the rest of the jais-family line, Aya Expanse and Command R7B Arabic. Announced but not obtainable: Falcon Arabic, Fanar Star, ALLaM 34B. Commercially usable is a fourth axis, crossing the other three. The gating labels here are the Hugging Face API's own gated field, read on 30 August 2026.
| Model (sizes) | Maker | Licence tag on the card | Weights downloadable |
|---|---|---|---|
| Jais 2 (8B, 70B) | Inception, MBZUAI, Cerebras | apache-2.0 | Yes, behind automatic gating |
| jais-family and jais-adapted (590M–70B) | Inception | apache-2.0 | Mixed: most auto-gated, a few ungated |
| ALLaM (7B, 13B, 70B in the paper; 34B in service) | National Center for AI, SDAIA | apache-2.0 in the frontmatter; the body points to a LICENSE file the repository does not contain | 7B only |
| Fanar 1 (9.24B base / 8.78B instruct) and Fanar 2 (27B) | QCRI, HBKU | apache-2.0 | Yes; Fanar Star (7B) not published |
| Falcon Arabic (7B), Falcon-H1 Arabic (3B, 7B, 34B) | TII | TII Falcon License, December 2024 | No public repository found |
| Karnak (6B, 40B) | Applied Innovation Center, Egypt | apache-2.0 | Yes |
| Aya Expanse (32B), Command R7B Arabic (8.03B) | Cohere Labs | cc-by-nc-4.0 | Yes, non-commercial |
The licence tag on a derivative is not the licence on the weights it started from
Fanar-1-9B-Instruct is a continual pretrain of google/gemma-2-9b and its card reads license: apache-2.0. Fanar-2-27B-Instruct, built on gemma-3-27b-pt, reads apache-2.0. AceGPT-v2-8B-Chat, a Llama-3 derivative, reads apache-2.0, as do the jais-adapted checkpoints. SILMA-9B, Atlas-Chat and Nile-Chat-12B start from the same Gemma bases and tag license: gemma.
Those documents do not say the same thing. Apache 2.0 imposes no use restriction, no attribution sentence, no naming rule and no user ceiling. The Gemma Terms of Use require distributors to "provide all third party recipients of Gemma or Model Derivatives a copy of this Agreement", forbid the uses in a separately hosted Prohibited Use Policy incorporated by reference, and reserve Google's right to "restrict (remotely or otherwise) usage" it believes violates the agreement. Meta's Llama 3 licence requires the phrase "Built with Meta Llama 3" displayed prominently, requires "Llama 3" at the start of a derivative's name, and requires a separate licence from Meta if you already had more than 700 million monthly active users on the version's release date.
Whether a particular relabelling is permissible is not something this article can settle; private agreements exist and are not published. What is checkable is that the two documents differ, that a tag has no authority over weights it did not create, and that other teams on the identical base chose the other label. ALLaM's own case is unresolved: the frontmatter says apache-2.0, the body says "Please see the LICENSE file", and the repository has none.
Falcon is widely described as Apache 2.0 and is not. Falcon 3 and Falcon-H1 point at the TII Falcon License of December 2024, royalty-free and permitting redistribution, but requiring a fixed sentence in public statements about a derivative, "[name of relevant Derivative Work] is built using artificial intelligence technology from the Technology Innovation Institute", and binding you to an acceptable-use policy the licensor may update after you ship. Aya Expanse and Command R7B Arabic are CC-BY-NC-4.0, non-commercial outright; UBC-NLP's NileChat-3B is research-only under a Qwen licence. Karnak-40B is the clean case: its base, Qwen3-30B-A3B-Instruct-2507, really is Apache 2.0, and so is Karnak-6B's Qwen3-4B-Instruct-2507. The name is not uniformly clean, though. Publicly listed GGUF quantisations declare Applied-Innovation-Center/Karnak-70B-LLAMA-v1.0 as their base model and carry license: llama3, and that source repository is not among the three models the organisation lists publicly.
Arabic-specialised does not mean ahead
QCRI's own table for Fanar-2-27B-Instruct records 69.40 on OALL v2 against 70.95 for google/gemma-3-27b-it. That is not the checkpoint it started from: the card's frontmatter names google/gemma-3-27b-pt, and the body says the team continually pretrained that model on approximately 166 billion tokens. The -it sibling is the instruction-tuned peer of an instruction-tuned model, so it is the like-for-like comparison rather than a softer one, and the table carries AceGPT-v2-32B-Chat and Qwen3-32B as further rows. On the same table Fanar leads ArabicMMLU 74.67 to 72.21 and Almieyar 79.46 to 70.48, and trails on GSM8K 93.70 to 95.80. One lab's runs, pointing in several directions at once.
Inception's Jais 2 card shows the trade running the other way. Jais 2 was trained from scratch, so there is no base model here and Llama-3.3-70B-Instruct is a peer rather than a predecessor: on IFEval, Jais-2-70B reports 74.53 Arabic instruction-level accuracy against Llama-3.3-70B-Instruct's 63.13; on the English half of the same test it reports 78.93 against Llama-3.3's 92.10. Every figure here is vendor-published, and nothing was re-run.
The most useful single result sits on the Fanar-1 card, because it holds the task fixed and varies only the register. On AraDiCE PIQA, Fanar-1-9B-Instruct scores 67.68 in Modern Standard Arabic, 63.66 in Egyptian, 59.03 in Levantine. ALLaM-7B-Instruct-preview: 67.52, 63.44, 60.88. AceGPT-v2-8B-Chat: 64.58, 61.32, 56.91. Llama-3.1-8B-Instruct, not Arabic-specialised at all: 58.11, 55.39, 54.24.
The benchmarks disagree about which Arabic they are measuring
ArabicMMLU (arXiv:2402.12840, MBZUAI, ACL Findings 2024) is 14,575 multiple-choice questions over 40 tasks, from school exams in Morocco, Egypt, Jordan, Palestine, Lebanon, the UAE, Kuwait and Saudi Arabia. Eight countries, but not eight comparable shares: in the dataset's own country field, the 14,455-item test split gives Jordan 5,990 questions, Egypt 2,487 and Palestine 2,032, roughly seven in ten of the split between them, against 101 for Saudi Arabia, with about 3,000 items carrying no country label at all. Every question is in Modern Standard Arabic, so this is curriculum diversity, not dialect diversity. The leaderboard's dialect signal comes down largely to Belebele (arXiv:2308.16884), of which OALL v2 keeps two Arabic tasks, MSA and Dialects; every Belebele question is built on a passage from FLORES-200, so the dialect versions are the same passages rendered into each variant.
AraDiCE (arXiv:2409.11404, QCRI, COLING 2025) describes its method plainly: machine translation followed by human post-editing, roughly 45,000 post-edited samples, covering Gulf, Egypt and the Levant. DialectalArabicMMLU (arXiv:2510.27543) adapts 3,000 MMLU-Redux items into Syrian, Egyptian, Emirati, Saudi and Moroccan. OALL v2 dropped v1 tasks in order to "remove saturated and machine translated tasks, due to inherent lower quality and possible cultural bias", and dialect coverage is where translated material survived anyway.
The suite a vendor picks decides what its number means. The evaluations directory inside the ALLaM repository names its Arabic tasks: acva, ar_ifeval, araMath_v3, araPro, arabicmmlu, etec_v2, exams_ar, gat, moe_ien_mcq, moe_ien_tf, openaimmlu. ETEC is Saudi Arabia's Education and Training Evaluation Commission, GAT the Saudi General Aptitude Test, moe_ien the Saudi education ministry's platform. ALLaM's strongest reported figure, 91.77 on IEN-MCQ, is a score on a Saudi national question bank.
TII's survey of the field (arXiv:2510.13430) catalogues more than 40 Arabic benchmarks and lists "low-resource dialect coverage (Sudanese, Mauritanian)" among its critical gaps. And the Gulf-Levant-Egypt scheme nearly every model card uses is contested in turn. Keleg, Goldwater and Magdy (arXiv:2505.21816, ACL 2025) extend a multi-label dataset in which speakers of eleven country-level dialects judge each sentence for their own variety, and report that four widely adopted assumptions, the first being that Arabic dialects can be grouped into distinguishable regional dialects, "oversimplify reality, and some of them are not always accurate", which they suggest may be hindering progress in tasks including dialect identification.
The instruments are small, defective in places, and built by interested parties
TII published QIMMA in April 2026 (arXiv:2604.03395), and it is a leaderboard and a benchmark audit at once: it consolidates 109 subsets and more than 52,000 samples, and screens every sample against a ten-point rubric before any model is scored on it. Reported discard rates run from 0.01% on GAT to 3.08% on ArabicMMLU, and inside ArabicMMLU they vary sharply by subject, with University Accounting at 83.8%. The failures are mundane: gold indices pointing at non-existent options, answer fields holding text that matches none of the listed choices, encoding corruption.
Two findings matter beyond bookkeeping. On the Arabic adaptations of the code benchmarks, 88% of HumanEval+ prompts and 81% of MBPP+ prompts required modification, which the authors attribute to "the state of the original Arabic translations rather than the difficulty of the task itself"; the changes they enumerate run from broken triple-quoted strings and indentation errors to normalisation toward more idiomatic Modern Standard Arabic. And on cultural benchmarks they report items framing "essentialist generalizations about specific populations as objectively correct answers", most evident in ArabCulture.
QIMMA is TII's work, and TII builds Falcon, though no Falcon model appears anywhere in QIMMA's tables. The clearer case is AraGen, whose 3C3H protocol several of these models report against: it is run by Inception, MBZUAI and Hugging Face, and two of those three build Jais. The AraGen-12-24 cycle was 279 questions judged by Claude 3.5 Sonnet, and 3C3H is a weighted composite, Correctness and Completeness scored 0 or 1 and Conciseness, Helpfulness, Honesty and Harmlessness graded 1 to 5 and normalised, so the headline figure is an average of six judgements by one model over 279 items. None of this is misconduct. It is a field small enough that the auditors and the audited are the same institutions.
The leaderboards lag as well. The OALL v2 results dataset was last modified on 3 March 2026 while its request queue was still taking submissions on 6 June, so a mid-2026 citation of "the leaderboard" may be a March snapshot. And Falcon's OALL figures are TII's own runs of the harness under TII's own namespace; among the 919 datasets in the OALL organisation there is no third-party result artefact for Falcon Arabic.
The short version
Four things decide a choice here and none is the headline score: whether the weights can actually be downloaded, what the base model's licence says rather than the derivative's tag, whether the benchmark tested Modern Standard Arabic or the dialect your users write, and how many items it contained.
More like this