How a Language Model Is Actually Built
A practical guide to training a language model from scratch, turning it into a chatbot, and why the only real limit is compute.
In short. There is no alchemy in a language model. It is built from four parts: a tokenizer that turns text into numbers, a corpus turned into one long stream of those numbers, a training loop that predicts the next number and corrects itself, and checkpoints so a long run survives interruption. Two more stages of the same kind turn it into a chatbot. The method is public, the code fits in a few hundred lines of Python, the data is free to download, and the training can run on your own PC overnight while you sleep. I trained nine models this way on one machine. What separates that from the models behind ChatGPT is scale: roughly a hundred million times more compute. And for a model that does a job for your organisation, a large open model fine-tuned on your own data is within reach of a single desk-side machine. This guide shows each step with working code, and where the real limit sits.
I get asked some version of the same question most weeks: how does AI actually work? Often it comes with something the person has read along the way. Nobody really knows how these models work. Building one takes a kind of knowledge only a handful of labs in the world have. It is closer to alchemy than engineering.
I understand where that comes from. The models are genuinely impressive, and most of what gets written about them stops at the impressive part. The answer to the question is far more down to earth, and I can show it, because I have built some.
This summer I trained nine small language models from scratch on a single machine, for the experiment in Testing Basque First on a Single Box. So here is the recipe, step by step, with the code. By the end you will know how a model like the one behind ChatGPT is built, what it would take to build your own, and where the real difference between my models and theirs lies.
There is no alchemy.
Start with the claim itself, because it bundles two different things together: how a model is built, and how a finished model arrives at a particular answer. The first is ordinary, published engineering.
The architecture most large models use, the transformer, was published in 2017. The training objective, predict the next token, is one line of code. The method that adjusts the weights, gradient descent with backpropagation, has been standard since the 1980s, and the optimiser most people use, AdamW, ships with PyTorch. Large, cleaned, openly licensed text corpora are a download away.
Open up a trained model and you find numbers. Mine holds 110 million of them, arranged in grids. Running it means multiplying and adding those grids. Training it means nudging each number, over and over, in whichever direction makes the next-token prediction slightly better.
There is a true part to “nobody knows how it works”. Once training ends, working out what any single number contributes, and why the model gives one answer and not another, is a live research field called interpretability. That is a question about reading the finished model. How to build one is fully known, and you can do it yourself.
What the big labs have is scale: more parameters, more data, and the compute to bring them together. That is a real barrier, and it is a barrier of money and electricity.
What you need.
- A GPU, ideally. I used a single NVIDIA DGX Spark (GB10 chip, 128GB of unified memory). Each of my nine runs took 17 to 20 hours. A gaming GPU runs the same code with a smaller batch or a smaller model, more slowly.
- Python and five libraries:
torch,transformers,tokenizers,datasetsandnumpy. - Text. I used FineWeb for English and FineWeb2 for Basque, both from Hugging Face under the ODC-BY 1.0 licence, which allows this use with attribution.
- Disk. A few gigabytes for the token streams, plus room for checkpoints.
Step one, the tokenizer.
A model cannot read letters. It reads numbers. A tokenizer splits text into pieces from a fixed vocabulary and gives each piece a number. Byte-pair encoding builds that vocabulary from your own text: it starts with single bytes and repeatedly merges the most common neighbouring pair into a new piece, until the vocabulary reaches the size you ask for.
from tokenizers import ByteLevelBPETokenizer
bpe = ByteLevelBPETokenizer()
bpe.train(
files=["sample.txt"], # a few hundred MB of your text, one document per line
vocab_size=32000,
min_frequency=2,
special_tokens=["<|endoftext|>"],
)
bpe.save_model("tokenizer") # writes vocab.json and merges.txtTrain it on a sample that looks like what the model will read. Mine was an equal mix of English and Basque, so neither language was split more clumsily than the other.
Step two, the corpus as a token stream.
Next, stream the corpus, turn every document into token numbers, and write them to disk as one long array. A vocabulary of 32,000 fits in 16 bits, so each token takes two bytes.
import numpy as np
from datasets import load_dataset
from transformers import GPT2TokenizerFast
tok = GPT2TokenizerFast(
vocab_file="tokenizer/vocab.json",
merges_file="tokenizer/merges.txt",
eos_token="<|endoftext|>",
)
ds = load_dataset("HuggingFaceFW/fineweb", name="sample-10BT",
split="train", streaming=True)
target, written = 242_000_000, 0
with open("train.bin", "ab") as f:
for row in ds:
ids = tok.encode(row["text"]) + [tok.eos_token_id]
np.array(ids, dtype=np.uint16).tofile(f)
written += len(ids)
if written >= target:
breakBefore training, set aside a slice the model never sees. Mine was 20 million tokens per language. It is the only honest way to tell learning from memorising.
Step three, the model and the training loop.
The model is a GPT-2-style decoder: 12 layers, 12 attention heads, 768 dimensions, a 1,024-token window, 110 million parameters. The transformers library builds it from a configuration in one call, with random weights.
The training loop is a single game, played tens of thousands of times. Take a batch of text, have the model predict each next token, measure how wrong it was, and nudge every weight to be slightly less wrong.
import numpy as np
import torch
import torch.nn.functional as F
from transformers import GPT2Config, GPT2LMHeadModel
data = np.memmap("train.bin", dtype=np.uint16, mode="r")
model = GPT2LMHeadModel(GPT2Config(
vocab_size=32000, n_positions=1024, n_embd=768, n_layer=12, n_head=12,
)).cuda()
opt = torch.optim.AdamW(model.parameters(), lr=3e-4,
betas=(0.9, 0.95), weight_decay=0.1)
seq_len, batch_size = 1024, 16
for step in range(start_step, total_steps):
starts = np.random.randint(0, len(data) - seq_len - 1, batch_size)
x = torch.stack([torch.from_numpy(data[i:i + seq_len].astype(np.int64))
for i in starts]).cuda()
y = torch.stack([torch.from_numpy(data[i + 1:i + 1 + seq_len].astype(np.int64))
for i in starts]).cuda()
logits = model(input_ids=x).logits
loss = F.cross_entropy(logits.reshape(-1, logits.size(-1)), y.reshape(-1))
opt.zero_grad()
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
opt.step()
if step % 500 == 0:
save_checkpoint(step) # see step fourThat is the whole of learning. y is x shifted along by one token, so every position is asked to predict the token that really came next. The loss measures the miss. backward works out how each of the 110 million weights contributed to it, and opt.step moves each one a little. My real loop also warms the learning rate up over the first 200 steps and logs progress, but nothing in it is more complicated than this.
At 16 sequences of 1,024 tokens, each step reads 16,384 tokens. My 484 million tokens per model took about 29,500 steps.
Step four, checkpoints.
A long run will be interrupted at some point. Save everything needed to carry on: the weights, the optimiser's running estimates, and the step reached. Write to a temporary file and rename it, so a crash mid-save never leaves a broken file.
import os
def save_checkpoint(step):
torch.save({"step": step,
"model": model.state_dict(),
"optimiser": opt.state_dict()}, "resume.tmp")
os.replace("resume.tmp", "resume.pt")On restart, load the file, restore the model and optimiser, and set start_step to the saved step. This one habit is what makes the next section possible.
Training while you sleep.
Because a run can stop and resume at any checkpoint, it does not need to run in one go. A PC that sits idle every night can train a model a few hours at a time, over as many nights as it takes.
One clarification on what does the work. Training is arithmetic, and the arithmetic runs on the GPU, or on the CPU if there is no GPU. Memory holds the model and the data while that happens. Spare memory on its own does not make training faster; idle processor time does.
To run it overnight:
Schedule it.
0 23 * * * cd /home/me/llm && timeout 8h python train.pyLet it resume.
resume.pt and carries on. If it is stopped mid-way, you lose at most the steps since the last checkpoint.Plan the number of nights.
Size the model to the machine.
Leave the case well ventilated. A GPU at full load for eight hours runs hot and draws real power, and that electricity is the cost of training.
Measuring it.
Perplexity on the held-out slice tells you how surprised the model is by text it has never seen. Average the loss over the held-out tokens and take its exponential. A perplexity of 50 means that at each token the model is about as unsure as if it were choosing between 50 equally likely options. Lower is better, and watching it fall over a run is the clearest sign the model is learning.
What you have at this point.
A file of 110 million numbers, about 220MB in half precision. Give it the start of some text and it continues it. This is called a base model.
prompt = tok("The harbour at dawn was", return_tensors="pt").input_ids.cuda()
out = model.generate(prompt, max_new_tokens=60, do_sample=True,
temperature=0.8, top_k=50)
print(tok.decode(out[0]))generate runs the same prediction in a loop: pick a likely next token, add it to the text, predict again. Temperature controls how adventurous the picks are. That loop is the whole of how any language model writes, including the large ones.
A base model will not answer a question. Ask it one and it may well continue with more questions, because that is what text containing a question often looks like.
Turning it into a chatbot.
A chatbot is a base model with two more stages of training, using the same loop on different data.
Stage one: supervised fine-tuning. Pick a fixed layout for conversations, with markers for who is speaking, and add those markers to the tokenizer:
tok.add_special_tokens({"additional_special_tokens": ["<|user|>", "<|assistant|>"]})
model.resize_token_embeddings(len(tok))Then train on thousands of example conversations written in that layout:
<|user|>What is the capital of France?<|assistant|>Paris.<|endoftext|>The loop is the one from step three, with two changes: it starts from your trained base model, not random weights, and it only counts the loss on the assistant's replies, so the model learns to answer and not to imitate the user. Open conversation datasets on Hugging Face supply the examples; check each one's licence before you use it. After this stage the model answers in turn.
Stage two: preference tuning. Show the model pairs of answers to the same question, one judged better than the other, and train it to make the better one more likely. The original method trained a separate reward model and used reinforcement learning. A newer method, direct preference optimisation, does it with one more variation of the same loss and loop. This stage is where tone, helpfulness and refusals are shaped.
Stage three: the chat window. Wrap generate in a loop that keeps the conversation so far, adds each new message with the right marker, and stops at the end-of-text token. A web page around that is a chatbot.
If you want to see every stage in one place, Andrej Karpathy's open-source nanochat runs the full pipeline, tokenizer to web chat interface, in about four hours for about $100 of rented GPU time.
Be ready for the result. A 110-million-parameter chatbot will answer in fluent, confident sentences, and will often be wrong. Quality comes from the size of the base model and the amount it was trained on. The method is the same at every size.
Where the real limit is.
Training compute can be estimated as six times the number of parameters times the number of training tokens. Put three models side by side:
| Model | Parameters | Training tokens | Compute | What it took |
|---|---|---|---|---|
| One of my nine runs | 110 million | 484 million | about 3 × 10¹⁷ operations | One DGX Spark, 17 to 20 hours |
| GPT-2 small, reproduced in llm.c (2024) | 124 million | 10 billion | about 7 × 10¹⁸ operations | Eight A100 GPUs, 90 minutes, about $20 |
| Llama 3, largest model (2024) | 405 billion | about 15 trillion | 3.8 × 10²⁵ operations | Meta's training cluster |
The code in this guide would train any of them. The difference is the bill. The largest Llama 3 model used roughly a hundred million times the compute of one of my runs.
Even my runs were short of what a model this size wants. A common rule of thumb from scaling research is about twenty training tokens per parameter. I used about four, to fit nine runs into a week. So the honest summary is this: the method is open to anyone, and good results are bought with compute and data.
To put a number on it for my own machine, take Google's Gemma 4 E2B, a small open model released in April 2026 with 2.3 billion effective parameters. Google has not published how much text it was trained on; the previous generation, Gemma 3, used 2 trillion tokens for its 1-billion-parameter model and 4 trillion for its 4-billion one. Taking 2 to 4 trillion tokens, an equivalent trained from scratch needs around 3 to 6 × 10²² operations. At the speed my Spark ran the Basque models, that is roughly 180 to 365 years. If well-tuned code ran ten times faster, it is still 17 to 35 years. A cluster of a thousand data-centre GPUs would do it in a day or two. Same recipe, same code. The only difference is the compute.
Fine-tuning a large model on your own data.
If the goal is a model that does a job for your organisation, training from scratch is rarely the route. Someone else has already paid for the expensive part. Gemma 4, released by Google under the Apache 2.0 licence, comes in four sizes, and you can take any of them and fine-tune it on your own documents, emails, tickets or reports. That is stage one of section ten, on a far stronger starting point, and it is well within reach of one desk-side machine.
The technique that makes it fit is LoRA. Instead of updating all of a model's weights, you freeze them and train a small set of extra weights alongside, typically well under one per cent of the total. QLoRA goes further and stores the frozen weights in 4-bit form, cutting their memory by about three-quarters. The result is an adapter file of a few hundred megabytes that changes how the base model behaves.
What fits on one DGX Spark, which has 128GB of memory:
| Model | What it is | Frozen weights, 4-bit | Frozen weights, 16-bit | Fits on |
|---|---|---|---|---|
| Gemma 4 E2B | 2.3 billion effective parameters | about 3GB | about 10GB | One Spark, easily |
| Gemma 4 26B A4B | Mixture of experts: 25.2 billion parameters, 3.8 billion used per token | about 13GB | about 50GB | One Spark |
| Gemma 4 31B | Dense, 31 billion parameters | about 16GB | about 62GB | One Spark |
| Gemma 4 31B, every weight trained | Full fine-tune with the AdamW optimiser | not applicable | about 500GB in total | Four linked Sparks, at the limit |
Those figures are the weights alone. Leave headroom for the working memory training needs, which grows with the length of the examples. NVIDIA's own guidance is that one 128GB Spark can fine-tune models of up to 70 billion parameters.
How long it takes. With LoRA, each training token costs about four times the model's active parameters in operations: the forward pass, plus working back through the frozen weights. For a company dataset of 10 million tokens, around 7.5 million words:
- Gemma 4 31B: about 1.2 × 10¹⁸ operations. Roughly three days at the speed my Spark ran the Basque models, and well under a day with well-tuned code.
- Gemma 4 26B A4B: only 3.8 billion parameters work on each token, so about 1.5 × 10¹⁷ operations, a matter of hours.
Two or three passes over the data multiply these. They are estimates from the same rule of thumb, not measurements, and 4-bit training runs somewhat slower than the arithmetic suggests. Measure your own throughput in the first few minutes and plan from that, as in section seven.
Linking Sparks. Up to four DGX Sparks can be linked through their 200Gb/s ConnectX-7 ports. Linking mainly buys memory: two give 256GB and four give 512GB, enough to train every weight of a 31-billion-parameter model, or to fine-tune on very long examples. It also shares out the work, but the link is far slower than the connections inside a data-centre server, so expect less than double the speed from two. For most fine-tuning, one Spark with LoRA is enough.
What fine-tuning is for. Fine-tuning changes how a model behaves: your house style, your document formats, your terminology, the shape of a good answer to your kind of question. It is a poor way to teach a model facts, especially facts that change. For questions like “what does our policy say”, retrieval, where the model looks up the relevant document at the moment it answers, is usually the better tool, and the two combine well. And for organisations that cannot send their data to a cloud provider, every step here, the data, the training and the finished model, stays inside the building.
What I am building this winter.
The most useful small model is often not a language model at all. It is a small model trained on one narrow job, with data only you have.
Mine is planned for this winter, for the boat. Over a season, the instruments log wind, boat speed and heading. A small model trained on those logs can learn how she actually sails, which is rarely what the published polars say, and give a live target speed at every wind angle. The second stage corrects the downloaded wind forecast for the waters I sail, by learning where it is reliably wrong. Both will run offline on a Raspberry Pi 5 with an AI accelerator board, with no signal needed at sea. I will write it up when it works.
Common questions.
Can anyone build their own LLM?
Yes. The method, the code and the data are public. A small model can be trained from scratch on one modern GPU in under a day, or over several nights on a home PC. What limits size and quality is compute and data.
Is building an AI model a secret or poorly understood process?
No. Every step, from the tokenizer to the training loop to chat tuning, is published and runs on open-source tools. What is still being researched is interpretability: explaining what a trained model's individual weights do.
How do you turn a language model into a chatbot?
Train the base model further on example conversations in a fixed format (supervised fine-tuning), then on pairs of better and worse answers (preference tuning), and wrap generation in a loop that keeps the conversation history.
Can you train a language model on a home PC?
Yes, at small sizes. Save checkpoints regularly and a run can be split across nights, starting when the PC is idle and resuming from the last checkpoint the next night. The GPU or CPU does the work, so idle processor time is what counts.
How much does it cost to train a language model?
It depends on size. GPT-2 small (124 million parameters) has been reproduced for about $20 of cloud GPU time. The largest Llama 3 model used 3.8 × 10²⁵ operations on Meta's own cluster, roughly a hundred million times the compute of a 110-million-parameter model trained for a day.
Should a business train its own model from scratch?
Rarely. Fine-tuning an existing open-weights model on your own data is faster and cheaper, and usually gives a better result for a specific task.
Can you fine-tune a large model on a DGX Spark?
Yes. With LoRA or QLoRA, a single 128GB DGX Spark can fine-tune open models such as Gemma 4 31B or Gemma 4 26B A4B on an organisation's own data, typically in hours to a few days for a dataset of around ten million tokens. Up to four Sparks can be linked for more memory.
How we work.
We build things to understand them, and we write up what we find. If you want to talk through where a model of your own would or would not help, the first conversation is a thirty-minute call, no pitch, no deck. We talk through where you are.
Sources.
- basque-first repository: the full pretraining pipeline this guide simplifies (MIT licence)
- nanochat (Karpathy, 2025): the full pipeline from tokenizer to chat interface
- Attention Is All You Need (Vaswani et al., 2017)
- Scaling Laws for Neural Language Models (Kaplan et al., 2020): the six-times-parameters-times-tokens estimate
- Training Compute-Optimal Large Language Models (Hoffmann et al., 2022): the twenty-tokens-per-parameter guidance
- Training language models to follow instructions with human feedback (Ouyang et al., 2022): supervised fine-tuning and preference tuning with a reward model
- Direct Preference Optimization (Rafailov et al., 2023)
- The Llama 3 Herd of Models (Meta, 2024)
- Reproducing GPT-2 (124M) in llm.c in 90 minutes for $20 (Karpathy, 2024)
- FineWeb and FineWeb2 (ODC-BY 1.0)
- Gemma 3 Technical Report (Google, 2025): training tokens per model size
- Gemma 4 model overview (DeepInfra): model sizes and active parameters
- Gemma 4 E2B (ApX): effective and total parameters, licence
- NVIDIA DGX Spark: memory, fine-tuning guidance and linking up to four systems
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., 2021)
- QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al., 2023)
The AI Already Switched On
A field note on what a Copilot and agents check found in a small-business Microsoft 365 tenancy, and the settings that put it under control.
Testing Basque First on a Single Box
A field note on training nine small language models from scratch to see whether the order a model learns its languages in leaves a lasting mark. Mostly it does not, and why is the interesting part.
The Shadow AI You Cannot See
A field note on a shadow-AI detector that runs entirely on your own machine, and what it says about governing the AI nobody authorised.