Iraqi Arabic and Language Models
Ask a model for Iraqi Arabic and one of three things usually happens: it answers in the standard language, it produces a blend of dialects, or it produces one Iraqi variety without telling you which. The third case is the interesting one, because "Iraqi Arabic" is not a variety. It names at least two, and those two split again inside the country.
"Iraqi Arabic" is not one language
The oldest classification in the literature is named after a single word. Blanc (1964, Harvard University Press) divided the Arabic of Mesopotamia into gilit and qeltu dialects after the way each pronounces the first-person perfect of the verb "to say" — gilit against qeltu, "I said". As reported in Yaseen (2015), qeltu is spoken by Muslims north of Baghdad, Mosul and Tikrit among them, and gilit across the central and southern regions and into parts of Iran. Jastrow (1978) subdivides qeltu further into Tigris, Euphrates and Anatolian groups.
The boundary between the two runs roughly from Falluja on the Euphrates to Samarra on the Tigris (Holes 2007, as reported in Yaseen). But it was never only a line on a map. Historically the Christian and Jewish communities of Baghdad spoke qeltu varieties, deep inside gilit territory — Blanc's book is titled for that fact, communal dialects within one city. No place label at any resolution, country or city, encodes it.
The two ISO 639-3 codes are already finer than anything anyone types into a prompt, and still too coarse. acm is Mesopotamian Arabic, which Glottolog names Gilit Mesopotamian Arabic; ayp is North Mesopotamian Arabic, for which Glottolog lists Mesopotamian Qeltu Arabic and Moslawi among the alternate names. The codes therefore bind roughly to the two types. But the ISO cut is north against south while the linguistic cut is gilit against qeltu, and those are not the same cut.
The boundary also moves. Palva (1983), as reported in Yaseen, describes qeltu as a geographically recessive type, losing ground to gilit inside Iraq and to Turkish and Kurdish in Anatolia; Collin (2009), reporting communication with Abu-Haidar and Holes, records Tikrit as formerly qeltu-speaking and now majority gilit. Yaseen's own study in Mosul found [q] almost categorically preserved across gender, age and social class — it is the local shibboleth — while the vowel [ɔː], which carries no comparable social weight, recedes in favour of the supralocal [uː], younger speakers using it less than older ones. Levelling is partial, and any dataset label is a snapshot of a boundary that has been shifting for at least forty years.
The standard English reference makes the point without setting out to. Erwin's A Short Reference Grammar of Iraqi Arabic, reissued by Georgetown University Press in 2004, describes — in the publisher's own words — Iraqi Arabic as spoken by Muslims in Baghdad. The most cited grammar of the language is a grammar of one community's variety in one city, and says so on the cover.
The lines inside Iraq do not agree with each other
Albuarabi's 2021 dissertation at the University of Wisconsin–Milwaukee treats Iraqi Arabic as a cluster of subdialects showing systematic microvariation in negation, and divides it in two on that basis. The ma group — Baghdadi, Najafi, Moslawi — negates sentences with the free morpheme ma and predicates with mu. The ma-š group — Amarah, Nasiriyah, Basrawi — uses a discontinuous ma…-š and the predicate negator muš.
Look at what that grouping does. It puts Baghdadi and Najafi, both gilit, together with Moslawi, which is qeltu, on one side; the other side is southern gilit. Negation and the qaf reflex draw different lines through the same country. A single label can get at most one of them right, and choosing a different label does not fix it.
The existential behaves the same way. aku "there is" and maku "there is not" belong to the ma group; the ma-š group has makuš or mamiš. The one token every list of Iraqi words leads with is not uniform across Iraq in its negated form. Albuarabi also records a negator ʕib in the Marshland dialect, citing Ingham and Hassan, and declines to analyse it — the variation continues below the level any published resource follows.
The present progressive splits three ways. "The student is studying in the library" takes da- in Baghdadi, ga- in Najafi and across the ma-š group, and kə-, qi- or ʕi- in Moslawi. The phonology bundles with it: /q/ surfaces as [g] everywhere except Moslawi, where it stays [q]; /r/ can surface as [ʁ] in Moslawi; /a/ raises to /i/ word-medially there. One sentence from the dissertation carries the whole bundle — Moslawi qɪlɪt-u … ʔə-qdəʁ … bɪ-ʔə-l-sˤuq against Najafi gɪlɪt-ləh … ʔə-gdər … bɪ-ʔə-l-sˤug. The shibboleth is in the first word.
What there is to learn from
Masader is the largest public catalogue of Arabic language and speech datasets. Counting the Dialect field across all 1,161 records in the main branch, retrieved on 31 August 2026: 598 Modern Standard, 329 mixed, 57 Egypt, 22 Morocco, 21 Levant, and 3 Iraq. The three are a 2022 corpus of 1,170 annotated posts released under CC0, and two 2006 telephone-speech collections of about fifty hours each, held by the Linguistic Data Consortium at the University of Pennsylvania — transcripts behind a $200 fee, audio on request, 478 conversation sides from 474 speakers.
Two corrections, both of which weaken the headline. Iraq also appears as one label inside multi-dialect sets: counting those, 25 datasets list Iraq against Egypt's 71. And the catalogue lags — neither of the two most substantial Iraqi corpora published since 2024 appears in it at all. "Three datasets" is a claim about the catalogue, not about the world, and the 25-against-71 figure is the more defensible of the two comparisons.
The recent Iraqi-dialect resources are small. CIAD (2022), from the universities of Kerbala and Babylon and Southern Technical University, is 1,170 posts annotated by three Iraqi Arabic experts, classified at 78% accuracy with a support vector machine; its authors justify building it in one clause, that no such corpus existed. IQAD (2025), built at Iran University of Science and Technology and published in Middle Technical University's Journal of Techniques, is far larger: 53,146 unique samples, 78,582 unique tokens, collected over four months from locality-focused public pages on a social platform, with TF-IDF and a linear classifier reaching 74%.
IQAD's label scheme is where it gets interesting. Its three classes are Middle, Western and Southern, assigned by the geography of the page rather than by linguistic feature — and searching its full text for north, Mosul, qeltu, gilit or Nineveh returns nothing. The largest Iraqi dialect dataset yet built has no category for the qeltu-speaking north, and its three regions correspond to neither the qaf line nor the negation line. The 2024 Ghadeer speech corpus is the other recent addition: 210 Iraqi speakers, 105 female and 105 male, 15,626 samples of three to six seconds each, CC BY — but balanced between Arabic and English, and the accompanying article does not state that Iraqi regional varieties are labelled per sample. A corpus of Iraqi speakers is not the same thing as a corpus annotated for Iraqi varieties.
What the benchmarks measure
MADAR, built at CMU Qatar and NYU Abu Dhabi, is a travel-phrase corpus commissioned in parallel across 25 Arab city dialects. Two thousand sentences were translated into all 25 cities plus Modern Standard; a further ten thousand were translated for five cities only — Beirut, Cairo, Doha, Tunis and Rabat. Iraq's three cities are Baghdad, Mosul and Basra, and none of them is in the second tier: 2,000 sentences each against 12,000. Note what MADAR gets right, though. Separating Mosul from Baghdad is the qeltu/gilit line, and separating Baghdad from Basra is the negation line. A corpus built by dialectologists draws exactly the distinctions a country label erases.
Its 2019 shared task asked systems to identify which of 26 classes a single sentence came from. The best macro-F1 was 67.32, against a prior state of the art of 67.9, and the top five systems used non-neural methods with word and character features. That is two-thirds accuracy on single sentences from one domain, with parallel training data for every class — conditions far friendlier than any real use.
NADI 2020 moved from commissioned translation to found social media text: 30,957 short posts, 21 countries, 100 provinces. Iraq's share was 3,816 posts across 12 provinces, 12.33% of the set and second only to Egypt, so it was not under-represented by count. The best macro-F1 was 26.78 at country level and 6.39 at province level. Then an Iraqi participating team, Aliwy et al. of the University of Kufa, manually inspected the training split for Farsi contamination and the organisers published the distribution: 504 of 21,000 posts overall, 2.40%. For Iraq, 382 of 2,556 — 14.95%. For Egypt, 2 of 4,473 — 0.04%. Roughly one in seven posts in the Iraqi training portion of a flagship Arabic dialect benchmark was in Persian. That is the difference between having data labelled Iraqi and having Iraqi data.
NADI 2023 reports 87.27 on country identification and it is not the same task: 18 countries instead of 21, posts filtered by geolocation, and classes deliberately balanced at 1,000 train, 100 dev and 200 test per country. The two numbers are not a progress curve. NADI 2024 then made the task multi-label and held out two undisclosed dialects — Iraq and Morocco — to test generalisation. The winning system scored 50.57 overall and 45.21 on the Iraqi samples, its lowest region, against 68.54 on the Nile Basin. Iraq was held out by design, so a low score is partly the design; the finding is not that models are bad at Iraqi but that Iraqi is the variety chosen as the generalisation test. Worth noting the schema underneath, too: the standard region-country-city aggregation files Iraq under "Gulf", which is defensible for gilit — the southern dialects are described in the literature as akin to Najdi — and simply wrong for qeltu.
The absences are the checkable part
What a benchmark leaves out is easier to verify than what it finds. The Arabic Online Commentary dataset, and the ALDi dialectness metric built on its 127,835 sentences, label Modern Standard plus Egyptian, Gulf and Levantine — no Iraqi. NADI 2023's dialect-to-standard translation subtasks covered Egyptian, Emirati, Jordanian and Palestinian. AraDiCE, a 2024 dialect and cultural evaluation of roughly 45,000 post-edited samples, covers Gulf, Egypt and the Levant. MADAR's ten-thousand-sentence tier covers five cities, none Iraqi. A 2026 study of whether dialects can be steered inside a model works on Egyptian, Moroccan, Levantine and Gulf.
One benchmark does name Iraqi as its own track: the 2024 OSACT6 dialect-to-standard translation task, covering Gulf, Egyptian, Levantine, Iraqi and Maghrebi. Its segments were drawn from a Saudi broadcast corpus, machine-translated into the standard language, then reviewed by native speakers of each dialect who could accept, edit or skip. Every dialect cleared 500 reviewed segments except Iraqi, which reached 277 — leaving a test set of 77 sentences against Gulf's 586. Results are reported as a single aggregate across all five dialects, so there is no published per-dialect score for Iraqi. The one benchmark with an Iraqi track does not report an Iraqi number.
DialectalArabicMMLU (2025) puts a figure on the general gap: across 19 models, mean accuracy was 62.8% in English and 51.9% in Modern Standard, then 49.8, 48.9, 48.2, 46.6 and 45.0 for Emirati, Egyptian, Saudi, Syrian and Moroccan. Every dialect sits below the standard; the standard sits eleven points below English. Iraqi is not in the table. So the size of the Iraqi gap is not merely large — it is unmeasured. Reading the benchmark literature above, I could find no published evaluation of Iraqi Arabic at the gilit/qeltu level, or at any sub-national level at all; MADAR's Baghdad/Mosul/Basra split is the closest thing, and it is a corpus, not an evaluation of what a model generates.
Why the output drifts to the middle
Keleg, Goldwater and Magdy (2025), at the University of Edinburgh, took 978 dialectal sentences with geolocations spread across the fourteen most populated Arab countries and gave them to 33 annotators from eleven countries, three per country, including three Iraqi annotators. Each judged, for their own country's dialect only, whether a speaker of it could have written the sentence. Just 249 of the 978 — about a quarter — turned out to be single-label at country level, and 56% were judged valid in more than one region. Most Arabic sentences are not diagnostic of any one country. A model asked for Iraqi can therefore produce paragraph after paragraph that an Iraqi annotator would accept as possibly Iraqi and that contains nothing Iraqi at all. That is the mechanism behind the usual complaint: it is not wrong, it just is not Iraqi.
The same paper contains the sharpest number in this literature. Evaluating a curated list of regional lexical cues against that annotated data, only 7 of 120 Iraqi-distinctive cues matched anything at all. Those 7 matched 7 sentences, of which 6 were genuinely Iraqi-valid and all 6 exclusively Iraqi — precision .86 and distinctiveness .86, the highest distinctiveness of the five regional lists it was scored against, and second only to Levantine on precision. But 204 sentences in the set were Iraqi-valid, so recall is .03. Iraqi lexical markers are excellent evidence when they appear, and they appear in about three per cent of Iraqi sentences. That is why a useful test has to force the diagnostic environments rather than read the prose.
As for what is happening inside the model, the 2026 steering study reports that models often default to the standard language or produce hybrid outputs that indiscriminately blend dialects, that dialect-specific neurons concentrate in late, generation-facing layers, and that the standard is clearly separated internally from the spoken dialects while the dialects share representation with one another. Read as an explanation — and this is interpretation, not their result — a request for one dialect lands in a region the model does not partition finely, so the high-probability output is either the separated standard cluster or a blend. Whether models fall back to one particular well-resourced dialect is widely asserted; I could find no study measuring it. What is measured is the drift to the standard and the blending.
A two-minute test
Do not ask for a paragraph in Iraqi Arabic and read it. By the numbers above you will learn almost nothing: most sentences carry no country signal, and the markers that do carry one appear in a small minority of them. Instead ask for four specific sentences in one prompt — "I said to him that I could sell it", "the student is studying in the library", "there is no food in the fridge", "Ali did not study" — and read the answer as a variety identification rather than a pass or a fail. The forms below are from Albuarabi's glossed examples.
Three outcomes are worth distinguishing. If everything returns in standard morphology, the model drifted — the commonest result and the easiest to see. If the answers are internally consistent, say gilit with da- and bare ma and maku, you have Baghdadi, which is a defensible answer to "Iraqi" and also a choice made silently on your behalf. If markers arrive from different rows at once — maku beside ga-, or qilit beside da- — you have a blend that no described variety produces, which is the failure the steering literature reports and the one that is invisible unless you force these environments.
One caution, because word-list articles usually get it wrong. The future particle raḥ is not a test: it is used far outside Iraq, and only its Moslawi realisation ʁaḥ, which follows the /r/ to [ʁ] rule, tells you anything regional. The same applies in reverse to ma…-š, which is genuinely southern Iraqi and also common across dialects hundreds of kilometres away. Neither presence nor absence of a single form settles anything; the profile across all four does.
| Force this | Modern Standard | Iraqi, by variety | What drift looks like |
|---|---|---|---|
| "I said" (1sg perfect of "to say") | qultu / qult | gilit — Baghdadi, Najafi, southern; qilit — Moslawi (qeltu is Blanc's label for the type; qilit is the Moslawi form Albuarabi records) | Standard qult with standard morphology around it; the [g]/[q] reflex is the oldest boundary in the literature |
| Present progressive, "is studying" | No progressive particle; bare imperfect yadrus | da-ydrus — Baghdadi; ga-ydrus — Najafi and the southern group; kə-/qi-/ʕi- — Moslawi | A bare imperfect is the standard; a b- or ʕam- prefixed imperfect belongs to varieties outside Iraq and is documented for none of these |
| Existential negation, "there is no food" | laysa hunāka / lā yūjad | maku — Baghdadi, Najafi, Moslawi; makuš or mamiš — Amarah, Nasiriyah, Basrawi | An existential built on a fī stem rather than the aku stem is not Iraqi in any variety described in these sources |
| Sentential negation, "Ali did not study" | lam / lā + verb | ma dɪrəs — the Baghdadi–Najafi–Moslawi group; ma-dərəs-iš — the southern group | lam plus jussive is standard drift. ma…-š is not drift: it identifies a southern Iraqi variety, and is also shared well outside Iraq |
| Future, "winter will come" | sa- / sawfa + imperfect | raḥ — Najafi; ʁaḥ — Moslawi, by the /r/ to [ʁ] rule | Not a diagnostic. raḥ is used far beyond Iraq; only the ʁ realisation carries regional information |
What to ask someone who claims Iraqi support
Which variety. A system described as supporting Iraqi that cannot answer gilit or qeltu, Baghdadi or Basrawi or Moslawi, is describing a label rather than a capability. The distinction is not pedantry: the negation line and the qaf line run in different directions, so a system tuned on one is not thereby correct on the other.
Measured on what, and on how many sentences. The one public shared task with an Iraqi track evaluates on 77 sentences and publishes no Iraqi-specific score. If a claim rests on a multi-dialect benchmark instead, ask what the Iraqi portion contained — in the 2020 collection the honest answer was that about one post in seven was in Persian.
Judged by whom. Native annotators of the variety concerned, not translators and not speakers of a neighbouring dialect, and not a model grading another model's output in a language for which no reference evaluation exists. For Iraqi the question already has a published answer: OSACT6 asked its reviewers for at least 500 segments per dialect, and Iraqi was the only one of the five that never reached that floor, stopping at 277.
And whether the evaluation is published at all. Since I could find no benchmark that evaluates Iraqi generation, any organisation making the claim — MentronX included — is making one it has to show its own working for, in public, in a form someone else can rerun. Until such an evaluation exists, the useful reply to any Iraqi-Arabic claim is three short questions: which variety, how many sentences, and who judged them.
The short version
Iraqi Arabic names at least two dialect types, and they split again on negation, so no single label describes it. Three Iraqi-primary datasets are catalogued; the one benchmark with an Iraqi track evaluates on 77 sentences. Force the diagnostic environments and read which variety you were handed.
More like this