Accepted for camera-ready · ICALP 2026De-romanizationArabizi
Breaking the Script Barrier
Automatic De-Romanization of Maghribi Arabizi to Arabic Script in Social Media
An empirical benchmark comparing an ordered rule mapper, MLE word mapping with regex fallback, a word-level character LSTM, and a full-passage mT5-small Transformer on the same 805-passage test set. Native-speaker audits and linguistic error analysis examine reference quality, alignment noise, and dialectal variation.
No single system leads on every metric. mT5 has the highest observed BLEU; MLE has the strongest character- and word-level form fidelity.
Accepted for camera-ready at the 9th International Conference on Arabic Language Processing. The revised manuscript has been submitted; final confirmation is pending. The full manuscript will be shared after confirmation. Code and aggregate-artifact access is available on request.
Four-system benchmark
PosEM measures position-wise token agreement. Bold values mark the best observed score in each column. The mT5–MLE BLEU difference is 1.31 points (95% paired-bootstrap interval: −0.55 to 3.24; p = 0.074), which does not establish superiority at the 0.05 level or equivalence.
Dataset, human audits & interpretation
Dataset: 8,045 passage pairs from UBC-NLP / NileChat Arabizi-Morocco, split into 6,436 training, 804 validation, and 805 test pairs. Its Arabizi was generated from Arabic-script Moroccan material with Command R+; performance on naturally authored Arabizi remains untested. Source data is subject to upstream access conditions and CC BY-NC 4.0.
Human audits: two native Moroccan Arabic speakers independently accepted all 300 sampled test references. All-positive labels make Cohen’s κ undefined. A separate 100-pair audit found 75 aligned and 25 misaligned positional training pairs. Review of 50 apparent LSTM errors yielded 42 true errors, 5 acceptable variants/non-errors, and 3 uncertain cases; this error-conditioned sample is not an accuracy estimate.
Interpretation: mT5 uses full passages, external pretraining, and more eligible training data than the word-level systems. This is not an architecture-controlled comparison. All 78 mT5 outputs that failed to emit EOS within the generation limit remain in the primary scores. Split provenance and source-document independence retain limitations described in the manuscript. The exploratory human mT5–MLE review does not establish a human-preference winner.