Testing Basque First on a Single Box
A field note on training nine small language models from scratch to see whether the order a model learns its languages in leaves a lasting mark.
David Eichler's paper argues something worth testing: that the language a model learns first sets its internal wiring, and that starting with a case-heavy language like Basque might give it a better grip on who did what to whom than starting with English ever could. Order first, everything else after.
Why Basque in particular? Because Basque marks who did what with word endings, not word order. English leans on position: the thing before the verb is usually the one doing the action. Basque instead stamps a suffix on the doer and lets the words fall where they like. If that kind of explicit grammar builds a sturdier internal sense of roles, Basque is where you would expect to see it.
It is a good idea, and an unusually hard one to test. So I spent a week actually testing it, on real hardware, and I want to report what I found, including the parts that failed, because the full picture is the useful one.
The catch: you cannot test it on a model you download.
Order is fixed the moment pretraining ends. Llama, Gemma, Latxa, every model you can download already has its training order baked in and frozen. Probe one and you learn about that model, not about order. That single fact rules out most of the takes I saw.
So I ran two things. One on shipped models, asking a smaller question I could actually answer. One from scratch, where I could set the order myself and watch what it did.
The rig.
Everything ran on one machine. A single NVIDIA DGX Spark: GB10 chip, 128GB of unified memory, CUDA 13, PyTorch 2.13. Both stages, all thirteen models, the same box.
One test does the heavy lifting in both stages, so let me be exact about it, with a real example rather than a stand-in. This is an actual sentence from the UD Basque-BDT treebank, a hand-checked collection of native Basque:
Lehenengo bi partzialetan huts batzuk egin ondoren, gero Kubak aise menderatu zuen Errusia.
It means “after some mistakes in the first two quarters, Cuba then comfortably defeated Russia.” The model is shown the Basque and made to choose between two English readings: “Cuba defeated Russia” or “Russia defeated Cuba.” Same words, same two countries, only the roles swapped. In Basque the answer is not in the word order, it is in the suffix: the -k on “Kubak” marks Cuba as the doer. Read the case and you get it right. Lean on word order and you can be walked straight off a cliff.
This kind of targeted test has a name: a probe. It is a small, narrow check for one specific ability, in this case whether the model tracks grammatical roles rather than guessing from position. I score it by likelihood, not by asking it a question. I feed the model the Basque sentence, measure how probable it finds each English reading token by token, average that, and whichever reading it rates as more likely is its answer. No prompting, no interpreting free text. Because the two options are identical except for who did what, nothing but the grammar can move the score.
The same scoring code runs on an 8-billion-parameter model and on a 110-million-parameter model I trained an hour earlier, because every model here is loaded in the same standard format. One ruler, used on both stages. And every Basque item is native and unaltered, never machine-translated, because translating with an English model would smuggle the English word-order habit straight into the test I am trying to run. Where I quote a margin, its error bars come from resampling the items thousands of times and reading off the middle 95%; for model-versus-model gaps the resampling is paired, so both models are always compared on the very same sentences.
Stage 1: what Basque does to a shipped model.
Narrow, answerable question. Does bolting Basque onto an English model change how it reads Basque grammar, or just make it more fluent?
Four models. Meta's Llama-3.1-8B, in base and chat form, against HiTZ's Latxa-Llama-3.1-8B, in base and chat form. Latxa is that same Llama with a large dose of extra Basque training added on top. So the pair changes exactly one thing and holds the rest fixed, which is the whole reason it is worth doing.
The Cuba sentence above is one of twenty-two Basque test items: eleven where case and word order agree, and eleven where they fight, which is the sharp test. Twenty-two is a small set and I am not pretending otherwise, which is precisely why every item is bracketed by two controls. A floor control in English where nothing in the sentence settles the answer, so any model should score around a coin-flip. And a ceiling control in English where the answer is spelled out, so any competent model should score near perfect. The floor landed around a half and the ceiling at essentially 1.0, exactly as they should, which is what tells me the numbers in between are trustworthy rather than noise. Alongside those I ran the Latxa team's own Basque exam suites (EusExams, EusProficiency, EusReading, EusTrivia) to separate “genuinely better at Basque grammar” from “just more fluent in Basque,” plus a couple of side probes on verb agreement and negation that came back flat or maxed-out for every model, so they do not carry the story.
The results are mixed, and I am reporting them mixed. Fluency moved a lot: on EusProficiency the Basque-adapted model jumps from 0.18 to 0.48, so the adaptation clearly took. The grammar probe is subtler. On the canonical items, the adapted model climbs from 0.64 to a clean 1.0 across all eleven, a gap whose error bars clear zero. On the conflict items, the ones that actually pit case against word order, it only edges up, 0.64 to 0.73, and the error bars touch zero. So Basque adaptation plainly helps ordinary Basque, and on the hardest structural test it is suggestive but not established.
The caveat that matters most: this is not Basque First. Latxa learned English first and Basque second. Stage 1 measures what adaptation does to an already-English model. It says nothing about which language came first. To get at that, you have to build the models yourself.
Stage 2: actually changing the order.
To test order you train from scratch and change only that. So there are no famous names in this section. I built the models, nine of them, from random noise, with no pretrained weights involved anywhere.
How you build one
Here is how you build one, because this is the part people are actually curious about, and it is more approachable than it looks. There are four moving parts.
First, a tokenizer. A model cannot read letters, it reads numbers, so you need something that chops text into a fixed vocabulary of pieces and gives each piece a number. I trained one shared 32,000-piece vocabulary on an equal mix of English and Basque, once, before anything else, so that neither language gets carved up more clumsily than the other. That matters here: Basque packs a lot of meaning into word endings, and an English-dominated vocabulary would shred those endings into confetti.
Second, the data. Stream real web text down from two public corpora, run every document through that tokenizer, and write each language out as one long ribbon of token numbers on disk. No magic, just millions of words turned into a stream of integers.
Third, the training loop, which is where the learning happens. The model plays one game, billions of times: look at a chunk of text, predict the next token, check the real answer, and nudge its internal numbers to be slightly less wrong. That nudging is gradient descent, run by an optimiser called AdamW, with the size of each nudge (the learning rate) warmed up and then slowly decayed over the run. A feeder hands it 1024-token chunks, drawn from the two ribbons in whatever English-to-Basque mix the curriculum demands at that exact moment.
Fourth, checkpoints. Every so often it writes its entire state to disk: the weights, the optimiser's running estimates, and the precise place it had reached in the data, so that a week-long run survives a reboot or me dropping off the connection.
That is the whole machine, and none of it is exotic. What you end up with, nine times over, is a single file of about 110 million numbers. Those numbers are the weights, and that file, roughly 220MB in half-precision, is the model. Feed it Basque or English and it predicts what comes next. The hard part was never the training. It was keeping all nine runs identical to the byte except for the one thing under test.
For the record, the shape of each model: a GPT-2-style decoder, the same family as the well-known nanoGPT, 12 layers, 12 attention heads, 768-dimensional, a 1024-token memory, 110 million parameters, half precision. If those numbers mean nothing to you, the only one that matters is the last: 110 million adjustable weights, which is tiny, roughly a thousandth the size of the models behind ChatGPT. A deliberately small brain, so I could afford to train nine of them.
The corpus and the curricula
The corpus is real web text, not toy data. Basque from FineWeb2 (the eus_Latn split), English from FineWeb, both cleaned through the same filtering pipeline so the two languages are not treated differently. 242 million tokens per language, 484 million per model, with a separate 20-million-token-per-language held-out set (text set aside and never trained on) kept back for honest scoring.
Three curricula. A learns English then Basque. B learns Basque then English. C interleaves the two evenly the whole way through, and is the control: the neutral baseline that favours neither order, the thing A and B are measured against. A and B do not switch cold. The first ~44% of training is one language alone, then it flips to 90/10 in favour of the second, with a 10% trickle of the first language kept flowing. That 44% is not a guess. It is the exact figure that makes total exposure come out at a dead-even 50/50 despite the trickle, so any difference between conditions is about order, not about how much of each language a model saw.
Before a single run started, I wrote down what would count as a real effect and did not touch it afterwards. The rule: the gap between curricula has to beat the gap between random seeds of the same curriculum. A seed is just the random starting point, the same recipe with a different roll of the dice. I ran three seeds per curriculum precisely so I could measure that roll-of-the-dice noise and refuse to call anything smaller than it an effect.
Each run took 17 to 20 hours, nine of them back to back, about seven days in all, running through a persistent session so it survived me dropping the connection. At every checkpoint I logged two things. Perplexity on the held-out text, which is how surprised the model is by text it has never seen (lower is better; a perplexity of 50 means that at each word it is about as unsure as if choosing between 50 equally likely options). And the grammar probe from the rig, run against the current weights.
One honest number before the results. At 484 million tokens for a 110-million-parameter model, these are undertrained on purpose, about four tokens per parameter against a rule-of-thumb twenty. That is what makes it cheap enough to run nine of them. It also means some effects may simply sit below the noise. Hold that thought for the probe.
Three results
The grammar probe: nothing. Every gap between curricula is smaller than the gap between seeds of the same curriculum. By my own pre-registered rule, that is a clean null. But the honest reading is “too small to tell,” not “no effect.” Undertrained at this size, the probe cannot resolve a structural difference either way. Figure 2 shows it plainly: every curriculum sitting on top of every other, well inside the noise.

Perplexity: a big, clean, repeatable order effect. The three seeds of a curriculum land within about a point of each other, at most 1.2. The gaps between curricula run from 7 to 52 points. So the effect is real and dwarfs the noise. But the cause is boring. A ends its training on Basque and B ends on English, and each is simply best at whatever it saw last. The first language is partly forgotten even with the 10% trickle running. That is recency and forgetting, not a model rebuilt around its first language.
The control converges. C lands with English and Basque within a fifth of a point of each other, and the lowest average perplexity of the three. That matches Foroutan and colleagues, who at far larger scale (1.1B and 3B parameters, 100B tokens) find that curriculum order changes the training path but not the final result. My control agrees with them.
Figure 1 tells the whole story in one picture: A and B as mirror images, each collapsing on its final language while the other drifts back up, and C gliding down the middle with both languages together.

| Curriculum | English perplexity | Basque perplexity | Grammar probe |
|---|---|---|---|
| A, English first | 87.3 | 47.0 | 0.46 |
| B, Basque first | 46.7 | 99.3 | 0.33 |
| C, interleaved control | 54.1 | 54.3 | 0.46 |
Means over three seeds. The seed-to-seed spread is about a point per cell, 1.2 at most, so the perplexity gaps are real. The probe column does not separate at all.
Common questions.
Does the order a language model learns its languages in change the final model?
At small scale, mostly no. In a controlled from-scratch experiment (nine 110M-parameter models, three curriculum orders, three seeds), curriculum order produced large perplexity differences, but those were driven by recency and forgetting, the model favouring whatever it saw last, not by a lasting change in how it encodes grammar. A grammar probe showed no order effect at this scale.
Can you test curriculum order on a released model like Llama or Latxa?
No. Training order is frozen once pretraining ends, so any downloaded model already has its order locked in. Probing one tells you about that model, not about order. Testing order requires training models from scratch and varying only the order.
What is the Basque First hypothesis?
The idea, from David Eichler, that the language a model is pretrained on first shapes its internal representational structure, and that starting with a morphologically rich, case-marking language like Basque might build a stronger grip on grammatical roles than starting with English.
How do you train a small language model from scratch?
Four parts: train a tokenizer to turn text into numbers, stream and tokenize a corpus into a token stream on disk, run a training loop that repeatedly predicts the next token and nudges the weights (gradient descent with AdamW), and checkpoint the state so a long run survives interruption. The result is a file of weights, here about 110 million numbers, roughly 220MB, that predicts the next token.
What did the experiment find about Basque First?
Basque adaptation clearly improves ordinary Basque competence in a shipped model, and on the hardest grammar contrast it is suggestive but not established. Curriculum order moves small-scale perplexity through forgetting, not deep restructuring. The strong Basque First claim remains untested, because a 110M undertrained model is too small to show it either way; settling it needs a full-scale run.
Where it lands.
So here is the scorecard. Basque adaptation clearly helps a model's ordinary Basque, and on the hardest grammar test it is suggestive but not established. Curriculum order does move the numbers at small scale, but through plain forgetting, the model favouring whatever it saw last, not through any deeper reshaping. And the real Basque First claim, that starting with a rich language leaves a lasting structural mark, is neither confirmed nor killed here, because a 110-million-parameter model trained on half a billion tokens is too small and too undertrained to show it either way.
Which is the actual point. The interesting hypotheses in this area are the expensive ones. A week on a single box will tell you which claims are cheap to check, will kill the ones that turn out to be measurement artefacts, and will price the experiment that could settle it: a full-scale Basque-first run against a matched English-first one, same tokenizer, same tests. It cannot settle the big question on its own.
Thanks to HiTZ for Latxa and the EusExams suite, and to David Eichler for a paper genuinely worth arguing with. The code, data licences and results are on GitHub: https://github.com/JustinNarracott/basque-first
The AI Already Switched On
A field note on what a Copilot and agents check found in a small-business Microsoft 365 tenancy, and the settings that put it under control.
The Shadow AI You Cannot See
A field note on a shadow-AI detector that runs entirely on your own machine, and what it says about governing the AI nobody authorised.
The Shape of UK Public Spending
A companion to the Tax Policy Associates map of the 85 UK taxes, built from the same OBR data. An interactive map of every pound of public spending in 2024-25, arranged by what it buys.