Saudi Arabia made AI a compulsory subject. What that asks of the technology underneath
Saudi Arabia made AI compulsory for more than six million pupils. What that implies for Arabic delivery, assessment, teacher capability and data.
Saudi Arabia's Ministry of Education has made artificial intelligence a compulsory school subject from the 2025-26 academic year, reaching more than six million pupils — one of the largest deployments of technology education attempted anywhere. For a school group, a publisher or a ministry team, the syllabus is the easy part. The hard part is that a compulsory subject has to be delivered in Arabic first rather than in English with a translation layer over it, assessed in a subject with no past papers and no item statistics, taught by teachers who learned the material weeks ahead of their pupils, and recorded in systems that Gulf regulators ask specific questions about.
None of those four is unusual on its own. What makes them hard together is the deadline. A pilot picks its schools; a mandate takes all of them, on the same morning.
A mandate is a different problem from a pilot
Pilots select for enthusiasm: a teacher who already reads about the subject at weekends, a working computer room, a head who wants the visit. A mandate also takes the school with one over-subscribed lab and a teacher covering the new subject alongside two others. Plan for the median school, not the flagship — shared devices, thin bandwidth at 8am when every class logs in at once, and sessions that survive a dropped connection without losing a pupil's work.
The second change is bureaucratic and easy to underestimate. A compulsory subject has to be evidenced upward — to a group office, a directorate, a ministry — not merely observed in the classroom. That obligation lands on whatever system holds the record, and it lands in the first term.
- Every school is in scope on the same date, including the ones that would never have volunteered.
- Coverage has to be reported upward, so "we taught it" needs to be a query someone can run, not a conversation.
- Pupils who move mid-year carry a record in the new subject with them, and the receiving school has to be able to read it.
- Guardians ask what the subject is and what their child is doing in it, often before the school has finished deciding.
- The first cohort's results become the baseline every later cohort is judged against, whether or not that baseline is any good.
Arabic first, not Arabic afterwards
Arabic-first delivery
A platform built so Arabic is a first-class language of authoring, delivery, assessment and administration: right-to-left layout on every screen including the administrative ones, content authored natively in Arabic rather than machine-translated from English, correct rendering where a left-to-right fragment — a code sample, a formula, an English acronym — sits inside a right-to-left sentence, and Arabic in every generated artefact from a question paper to a certificate. A translation layer, by contrast, translates the interface strings and leaves everything else in the shape English gave it.
This subject punishes a translation layer harder than most. AI teaching material is full of left-to-right fragments — Python snippets, library names, URLs, notation — embedded in Arabic prose. Bidirectional rendering is where a platform fails visibly: brackets that reverse, punctuation that jumps to the wrong end of a line, a code block that reads correctly in the editor and wrongly in the printed worksheet. Ask to see a lesson with a code sample, in Arabic, on a phone. It takes a minute and settles the question.
Then check the side nobody demos. The learner interface is usually the translated one; the coordinator's screens, the grading queue and the exported PDF are where mirrored layouts get missed. A coordinator who switches to English to build a coverage report is doing the work twice, and will end up keeping the real record in a spreadsheet.
What does Arabic-first actually mean for a learning platform?
That Arabic is a language the product was built in, not one it was translated into. Concretely: right-to-left layout across learner, teacher and administrator screens rather than the learner ones only; mixed-direction text that renders correctly when a code sample or formula sits inside an Arabic sentence; Arabic in generated documents, certificates and exports, not only on screen; and search, sorting and question types that behave correctly with Arabic text. The quickest test in a demo is to ask for the administrative screens in Arabic on a small viewport — translation layers are usually complete on the learner side and thin everywhere else.
Content sourcing deserves its own decision. There is far more AI teaching material in English than in Arabic, so most groups will translate or generate. Machine translation across a whole textbook produces terminology drift — the same concept named three ways in three chapters, which is fatal in a subject where the vocabulary is the learning. Fix a glossary in Modern Standard Arabic first, agreed with the teachers who will use it, then hold every unit to it and have a subject teacher review the result, not a translator.
Assessing a subject nobody has taught before
Everything a mature subject leans on is missing in year one. No past papers, no anchor scripts, no cut scores, no item difficulty statistics, no examiner with fifteen years of moderation behind them. So the first year's assessment carries two jobs at once: report honestly on this cohort, and build the evidence base that makes the second year defensible.
Write outcomes as observable competencies before writing any questions. "Can explain why a model trained on unbalanced data misclassifies the smaller group" can be moderated across forty schools; "understands machine learning" cannot. In a new subject the outcome statement is the only fixed point you have, which is why competency frameworks earn their keep here more than in a settled one.
Use rubrics deliberately rather than by habit. Analytic when you want a diagnostic profile of where a cohort is weak — and in year one you badly want that. Holistic when you need defensible marks quickly at scale. Single-point when the target is a published standard. Then keep the item statistics from the first sitting: difficulty and discrimination per item, so the question everybody answered correctly and the question nobody did can be repaired before they are reused. Ordinary psychometrics matter more in year one than in year ten. The machinery for all of it — rubrics, moderation, item statistics — is set out on the assessments page.
One problem is specific to this subject: pupils will use a model to do the homework about models. Design around it rather than policing it — in-class performance tasks, portfolios that keep the drafts as well as the final piece, short oral defences of a project. Using the tool well is part of what you are teaching, so assess the judgement rather than the typing.
How do you assess a school subject with no prior data?
Define observable competencies first and treat them as the unit of record, because a competency can be moderated across schools when a mark cannot. Assess through performance tasks and portfolios rather than recall papers, since neither an item bank nor a mark scheme exists yet. Use rubrics — analytic where you want a diagnostic profile, holistic where you need marks at scale — and keep every rubric decision. Then capture item statistics from the first sitting so year two starts with calibrated questions, and say openly that year one's numbers are calibration rather than judgement.
Teacher capability is the constraint that binds
Do the arithmetic before buying anything: teachers to be trained, times hours each, divided by the weeks before term, against the people who can actually deliver the training. In most groups the number does not work, and the usual answer is cascade training — train the trainers, who train the trainers, who train the teachers. Fidelity drops at every hop. A fourth-hand lesson on neural networks is not a lesson on neural networks.
What survives the cascade better is material a teacher can hold: the same unit the pupil gets, worked answers, and somewhere to ask a question and get one back from someone who knows. Give teachers the same platform as the learners, so there is one record and one set of content rather than a training portal beside a learning portal with separate logins and no shared reporting.
Measure completion, not invitation. Attendance is not evidence that anybody can teach the material; a short assessment and a certificate naming what was demonstrated is closer. And build one channel back — which lessons collapse in a real classroom with thirty pupils and twelve working machines. That feedback is worth more than any inspection in year one, and it has a short shelf life: teachers stop reporting problems once they decide nobody reads the reports. A group running many campuses has more of this to coordinate than a single school does; the schools page covers the group-level version.
Data protection questions get asked differently in the Gulf
Saudi Arabia's Personal Data Protection Law, administered by SDAIA, sets conditions on how personal data is handled and on transferring it outside the Kingdom. Procurement here opens with residency, transfer and sub-processors, not with a feature list. Teaching the subject with AI tools adds a question a geography lesson never raised: what happens to a pupil's work when it is sent to a model, and whose model is it.
Get these answered in writing, before signature, from every vendor including us.
- Where is pupil data stored, in which country, and what is the notice period if that answer changes?
- Which sub-processors touch it — naming any model provider behind an AI feature — and where does each of them run?
- Is pupil work used to train anyone's model? "No" belongs in the contract, not in a blog post.
- How long is data kept after a pupil leaves, who can delete it, and is deletion demonstrable?
- What can a guardian see, and can that be scoped by category rather than all-or-nothing?
- Can you produce a record of who read or changed a pupil's data, and can that record be shown to be unaltered?
- Which of the vendor's own staff can see pupil data, under what approval, and is that approval logged?
"We are GDPR compliant" is not an answer to a PDPL question. It usually means the vendor has done the engineering that makes a PDPL answer possible — lawful-basis handling, consent records, subject access, erasure — but the mapping is your legal team's to check, not the sales engineer's to assert.
Does a Saudi school need its learner data stored inside the Kingdom?
That is a question for your legal advisers rather than a vendor, because it turns on the data, the purpose and the implementing regulations under the Personal Data Protection Law, which SDAIA administers. The buyer's practical position is simpler: ask every vendor to name the storage country, list the sub-processors including any model provider, and commit contractually to notifying you before either changes. A vendor who cannot name the country in the first meeting will not be faster in the third.
What a platform can and cannot take off you
No platform teaches this subject. What it removes is the duplicated work that eats a rollout: authoring the same unit once per school instead of once, chasing which sites have covered which outcome, exporting spreadsheets to answer one question from the directorate, and running one system for pupils and another for the teachers being trained to teach them.
One tenant with an organisation under it per school, one item bank, one competency framework, and coverage across every site answerable as a query.
A separate installation per school, a shared drive of documents, and a data request that takes three weeks and produces a number nobody quite trusts.
That is the shape of platform we build, so read what follows as interested rather than neutral. Lurno ships in four product languages — English, Arabic with right-to-left layout throughout, French and Italian — and the Arabic covers the administrative screens, not only the learner ones. On the assessment side: 19 question types, 7 assessment types, rubric grading in analytic, holistic and single-point forms, peer review, portfolios, a unified grading queue, and psychometrics including adaptive IRT/CAT for when an item bank is mature enough to justify it. Competency frameworks carry evidence and automatic mastery, which is what turns "we taught it" into something a directorate can inspect. Tenant isolation is enforced in the database rather than the application — separation enforced underneath the platform — and sensitive changes are written to a hash-chained audit log.
And plainly, the parts that are not there. SCORM 1.2, SCORM 2004, xAPI and LTI 1.3 are in development — modelled in the product, runtime still being built — so if a tender requires standards conformance on day one, say so at the first meeting. SAML/OIDC SSO, MFA and passkeys are on the roadmap; what exists today is partner sign-on (silent SSO), a different thing that should not be sold as the same one. Payments, checkout and self-serve signup are not built: accounts are created by invitation. We hold no SOC 2 or ISO 27001 certificate; the practices are aligned to ISO 27001 and the security page says which.
The groups that will look competent in three years are the ones treating this first year as calibration: outcomes written before questions, item statistics kept from the first sitting, teacher capability measured rather than assumed, Arabic tested on the screens nobody demos, and the data questions answered in the contract while the contract is still open. The syllabus will be revised — new subjects always are. The record of what was taught, to whom, and how well it was learned is what you will still need when it is.