NavitecnAvItec
PERSPECTIVES · NAVITEC
Field Note

How a Language Model Is Actually Built

A practical guide to training a language model from scratch, turning it into a chatbot, and why the only real limit is compute.

In short. There is no alchemy in a language model. It is built from four parts: a tokenizer that turns text into numbers, a corpus turned into one long stream of those numbers, a training loop that predicts the next number and corrects itself, and checkpoints so a long run survives interruption. Two more stages of the same kind turn it into a chatbot. The method is public, the code fits in a few hundred lines of Python, the data is free to download, and the training can run on your own PC overnight while you sleep. I trained nine models this way on one machine. What separates that from the models behind ChatGPT is scale: roughly a hundred million times more compute. And for a model that does a job for your organisation, a large open model fine-tuned on your own data is within reach of a single desk-side machine. This guide shows each step with working code, and where the real limit sits.

I get asked some version of the same question most weeks: how does AI actually work? Often it comes with something the person has read along the way. Nobody really knows how these models work. Building one takes a kind of knowledge only a handful of labs in the world have. It is closer to alchemy than engineering.

I understand where that comes from. The models are genuinely impressive, and most of what gets written about them stops at the impressive part. The answer to the question is far more down to earth, and I can show it, because I have built some.

This summer I trained nine small language models from scratch on a single machine, for the experiment in Testing Basque First on a Single Box. So here is the recipe, step by step, with the code. By the end you will know how a model like the one behind ChatGPT is built, what it would take to build your own, and where the real difference between my models and theirs lies.

SECTION ONE · NO ALCHEMY

There is no alchemy.

Start with the claim itself, because it bundles two different things together: how a model is built, and how a finished model arrives at a particular answer. The first is ordinary, published engineering.

The architecture most large models use, the transformer, was published in 2017. The training objective, predict the next token, is one line of code. The method that adjusts the weights, gradient descent with backpropagation, has been standard since the 1980s, and the optimiser most people use, AdamW, ships with PyTorch. Large, cleaned, openly licensed text corpora are a download away.

Open up a trained model and you find numbers. Mine holds 110 million of them, arranged in grids. Running it means multiplying and adding those grids. Training it means nudging each number, over and over, in whichever direction makes the next-token prediction slightly better.

There is a true part to “nobody knows how it works”. Once training ends, working out what any single number contributes, and why the model gives one answer and not another, is a live research field called interpretability. That is a question about reading the finished model. How to build one is fully known, and you can do it yourself.

What the big labs have is scale: more parameters, more data, and the compute to bring them together. That is a real barrier, and it is a barrier of money and electricity.

Diagram of the stages of building a language model: text corpus, tokenizer, token stream, training loop, base model, supervised fine-tuning, preference tuning, chatbot. THIS GUIDE, STEPS ONE TO FOUR SECTION TEN Text corpus Tokenizer Token stream Training loop Base model Supervised fine-tuning Preference tuning Chatbot 1 Predict the next token 2 Measure the miss 3 Nudge the weights
Diagram of the stages of building a language model: text corpus, tokenizer, token stream, training loop, base model, supervised fine-tuning, preference tuning, chatbot. THIS GUIDE, STEPS ONE TO FOUR Text corpus Tokenizer Token stream Training loop 1 Predict the next token 2 Measure the miss 3 Nudge the weights Base model SECTION TEN Supervised fine-tuning Preference tuning Chatbot
Figure 1. Every stage of building a language model, from raw text to a chatbot. Each one is published and runs on open-source tools.
SECTION TWO · WHAT YOU NEED

What you need.

  • A GPU, ideally. I used a single NVIDIA DGX Spark (GB10 chip, 128GB of unified memory). Each of my nine runs took 17 to 20 hours. A gaming GPU runs the same code with a smaller batch or a smaller model, more slowly.
  • Python and five libraries: torch, transformers, tokenizers, datasets and numpy.
  • Text. I used FineWeb for English and FineWeb2 for Basque, both from Hugging Face under the ODC-BY 1.0 licence, which allows this use with attribution.
  • Disk. A few gigabytes for the token streams, plus room for checkpoints.
SECTION THREE · THE TOKENIZER

Step one, the tokenizer.

A model cannot read letters. It reads numbers. A tokenizer splits text into pieces from a fixed vocabulary and gives each piece a number. Byte-pair encoding builds that vocabulary from your own text: it starts with single bytes and repeatedly merges the most common neighbouring pair into a new piece, until the vocabulary reaches the size you ask for.

from tokenizers import ByteLevelBPETokenizer

bpe = ByteLevelBPETokenizer()
bpe.train(
    files=["sample.txt"],        # a few hundred MB of your text, one document per line
    vocab_size=32000,
    min_frequency=2,
    special_tokens=["<|endoftext|>"],
)
bpe.save_model("tokenizer")      # writes vocab.json and merges.txt

Train it on a sample that looks like what the model will read. Mine was an equal mix of English and Basque, so neither language was split more clumsily than the other.

SECTION FOUR · THE TOKEN STREAM

Step two, the corpus as a token stream.

Next, stream the corpus, turn every document into token numbers, and write them to disk as one long array. A vocabulary of 32,000 fits in 16 bits, so each token takes two bytes.

import numpy as np
from datasets import load_dataset
from transformers import GPT2TokenizerFast

tok = GPT2TokenizerFast(
    vocab_file="tokenizer/vocab.json",
    merges_file="tokenizer/merges.txt",
    eos_token="<|endoftext|>",
)
ds = load_dataset("HuggingFaceFW/fineweb", name="sample-10BT",
                  split="train", streaming=True)

target, written = 242_000_000, 0
with open("train.bin", "ab") as f:
    for row in ds:
        ids = tok.encode(row["text"]) + [tok.eos_token_id]
        np.array(ids, dtype=np.uint16).tofile(f)
        written += len(ids)
        if written >= target:
            break

Before training, set aside a slice the model never sees. Mine was 20 million tokens per language. It is the only honest way to tell learning from memorising.

SECTION FIVE · THE TRAINING LOOP

Step three, the model and the training loop.

The model is a GPT-2-style decoder: 12 layers, 12 attention heads, 768 dimensions, a 1,024-token window, 110 million parameters. The transformers library builds it from a configuration in one call, with random weights.

The training loop is a single game, played tens of thousands of times. Take a batch of text, have the model predict each next token, measure how wrong it was, and nudge every weight to be slightly less wrong.

import numpy as np
import torch
import torch.nn.functional as F
from transformers import GPT2Config, GPT2LMHeadModel

data = np.memmap("train.bin", dtype=np.uint16, mode="r")
model = GPT2LMHeadModel(GPT2Config(
    vocab_size=32000, n_positions=1024, n_embd=768, n_layer=12, n_head=12,
)).cuda()
opt = torch.optim.AdamW(model.parameters(), lr=3e-4,
                        betas=(0.9, 0.95), weight_decay=0.1)

seq_len, batch_size = 1024, 16
for step in range(start_step, total_steps):
    starts = np.random.randint(0, len(data) - seq_len - 1, batch_size)
    x = torch.stack([torch.from_numpy(data[i:i + seq_len].astype(np.int64))
                     for i in starts]).cuda()
    y = torch.stack([torch.from_numpy(data[i + 1:i + 1 + seq_len].astype(np.int64))
                     for i in starts]).cuda()

    logits = model(input_ids=x).logits
    loss = F.cross_entropy(logits.reshape(-1, logits.size(-1)), y.reshape(-1))

    opt.zero_grad()
    loss.backward()
    torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
    opt.step()

    if step % 500 == 0:
        save_checkpoint(step)    # see step four

That is the whole of learning. y is x shifted along by one token, so every position is asked to predict the token that really came next. The loss measures the miss. backward works out how each of the 110 million weights contributed to it, and opt.step moves each one a little. My real loop also warms the learning rate up over the first 200 steps and logs progress, but nothing in it is more complicated than this.

At 16 sequences of 1,024 tokens, each step reads 16,384 tokens. My 484 million tokens per model took about 29,500 steps.

SECTION SIX · CHECKPOINTS

Step four, checkpoints.

A long run will be interrupted at some point. Save everything needed to carry on: the weights, the optimiser's running estimates, and the step reached. Write to a temporary file and rename it, so a crash mid-save never leaves a broken file.

import os

def save_checkpoint(step):
    torch.save({"step": step,
                "model": model.state_dict(),
                "optimiser": opt.state_dict()}, "resume.tmp")
    os.replace("resume.tmp", "resume.pt")

On restart, load the file, restore the model and optimiser, and set start_step to the saved step. This one habit is what makes the next section possible.

SECTION SEVEN · WHILE YOU SLEEP

Training while you sleep.

Because a run can stop and resume at any checkpoint, it does not need to run in one go. A PC that sits idle every night can train a model a few hours at a time, over as many nights as it takes.

One clarification on what does the work. Training is arithmetic, and the arithmetic runs on the GPU, or on the CPU if there is no GPU. Memory holds the model and the data while that happens. Spare memory on its own does not make training faster; idle processor time does.

To run it overnight:

1

Schedule it.

On Windows, use Task Scheduler to start the script at 11pm and stop it at 7am. On Linux or a Mac, a single cron line does both:
0 23 * * * cd /home/me/llm && timeout 8h python train.py
2

Let it resume.

Each night the script loads resume.pt and carries on. If it is stopped mid-way, you lose at most the steps since the last checkpoint.
3

Plan the number of nights.

Run for a few minutes, note the tokens per second, then divide your token target by tokens per second times 28,800 (the seconds in eight hours).
4

Size the model to the machine.

If the arithmetic says it would take months, shrink the model: fewer layers and a narrower width cut both the memory needed and the time per step. A model of a few million parameters will train on a laptop CPU, slowly.

Leave the case well ventilated. A GPU at full load for eight hours runs hot and draws real power, and that electricity is the cost of training.

SECTION EIGHT · MEASURING IT

Measuring it.

Perplexity on the held-out slice tells you how surprised the model is by text it has never seen. Average the loss over the held-out tokens and take its exponential. A perplexity of 50 means that at each token the model is about as unsure as if it were choosing between 50 equally likely options. Lower is better, and watching it fall over a run is the clearest sign the model is learning.

Line chart of held-out perplexity for English and Basque falling from about 143 and 153 to about 54.1 and 54.3 over 484 million training tokens. 50 60 80 100 150 0 100 200 300 400 500 Tokens trained (millions) Held-out perplexity (log scale) English 54.1 Basque 54.3
Line chart of held-out perplexity for English and Basque falling from about 143 and 153 to about 54.1 and 54.3 over 484 million training tokens. 50 100 150 0 200 400 Tokens trained (millions) Held-out perplexity (log scale) English 54.1 Basque 54.3
Figure 2. Held-out perplexity falling as one of my models trains: the interleaved run, mean of three seeds. Lower means the model is less surprised by text it has never seen.
SECTION NINE · THE BASE MODEL

What you have at this point.

A file of 110 million numbers, about 220MB in half precision. Give it the start of some text and it continues it. This is called a base model.

prompt = tok("The harbour at dawn was", return_tensors="pt").input_ids.cuda()
out = model.generate(prompt, max_new_tokens=60, do_sample=True,
                     temperature=0.8, top_k=50)
print(tok.decode(out[0]))

generate runs the same prediction in a loop: pick a likely next token, add it to the text, predict again. Temperature controls how adventurous the picks are. That loop is the whole of how any language model writes, including the large ones.

A base model will not answer a question. Ask it one and it may well continue with more questions, because that is what text containing a question often looks like.

SECTION TEN · THE CHATBOT

Turning it into a chatbot.

A chatbot is a base model with two more stages of training, using the same loop on different data.

Stage one: supervised fine-tuning. Pick a fixed layout for conversations, with markers for who is speaking, and add those markers to the tokenizer:

tok.add_special_tokens({"additional_special_tokens": ["<|user|>", "<|assistant|>"]})
model.resize_token_embeddings(len(tok))

Then train on thousands of example conversations written in that layout:

<|user|>What is the capital of France?<|assistant|>Paris.<|endoftext|>

The loop is the one from step three, with two changes: it starts from your trained base model, not random weights, and it only counts the loss on the assistant's replies, so the model learns to answer and not to imitate the user. Open conversation datasets on Hugging Face supply the examples; check each one's licence before you use it. After this stage the model answers in turn.

Stage two: preference tuning. Show the model pairs of answers to the same question, one judged better than the other, and train it to make the better one more likely. The original method trained a separate reward model and used reinforcement learning. A newer method, direct preference optimisation, does it with one more variation of the same loss and loop. This stage is where tone, helpfulness and refusals are shaped.

Stage three: the chat window. Wrap generate in a loop that keeps the conversation so far, adds each new message with the right marker, and stops at the end-of-text token. A web page around that is a chatbot.

If you want to see every stage in one place, Andrej Karpathy's open-source nanochat runs the full pipeline, tokenizer to web chat interface, in about four hours for about $100 of rented GPU time.

Be ready for the result. A 110-million-parameter chatbot will answer in fluent, confident sentences, and will often be wrong. Quality comes from the size of the base model and the amount it was trained on. The method is the same at every size.

SECTION ELEVEN · THE REAL LIMIT

Where the real limit is.

Training compute can be estimated as six times the number of parameters times the number of training tokens. Put three models side by side:

ModelParametersTraining tokensComputeWhat it took
One of my nine runs110 million484 millionabout 3 × 10¹⁷ operationsOne DGX Spark, 17 to 20 hours
GPT-2 small, reproduced in llm.c (2024)124 million10 billionabout 7 × 10¹⁸ operationsEight A100 GPUs, 90 minutes, about $20
Llama 3, largest model (2024)405 billionabout 15 trillion3.8 × 10²⁵ operationsMeta's training cluster
Bar chart on a log scale comparing training compute: one of my runs about 3 times 10 to the 17, GPT-2 small about 7 times 10 to the 18, a Gemma 4 E2B equivalent about 3 to 6 times 10 to the 22, and Llama 3 at 3.8 times 10 to the 25. One of my runs about 3 × 1017 GPT-2 small, llm.c about 7 × 1018 Gemma 4 E2B equivalent 3 to 6 × 1022 (estimate) Llama 3, largest 3.8 × 1025 about 100 million times 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 Training operations (log scale, each step ten times the last)
Bar chart on a log scale comparing training compute: one of my runs about 3 times 10 to the 17, GPT-2 small about 7 times 10 to the 18, a Gemma 4 E2B equivalent about 3 to 6 times 10 to the 22, and Llama 3 at 3.8 times 10 to the 25. One of my runs, about 3 × 1017 GPT-2 small, llm.c, about 7 × 1018 Gemma 4 E2B equivalent, 3 to 6 × 1022 (estimate) Llama 3, largest, 3.8 × 1025 about 100 million times 1016 1018 1020 1022 1024 1026 Training operations (log scale)
Figure 3. Training compute on a log scale. Each step on the axis is ten times the last. The Gemma 4 E2B bar is an estimate, because Google has not published its training data size.

The code in this guide would train any of them. The difference is the bill. The largest Llama 3 model used roughly a hundred million times the compute of one of my runs.

Even my runs were short of what a model this size wants. A common rule of thumb from scaling research is about twenty training tokens per parameter. I used about four, to fit nine runs into a week. So the honest summary is this: the method is open to anyone, and good results are bought with compute and data.

To put a number on it for my own machine, take Google's Gemma 4 E2B, a small open model released in April 2026 with 2.3 billion effective parameters. Google has not published how much text it was trained on; the previous generation, Gemma 3, used 2 trillion tokens for its 1-billion-parameter model and 4 trillion for its 4-billion one. Taking 2 to 4 trillion tokens, an equivalent trained from scratch needs around 3 to 6 × 10²² operations. At the speed my Spark ran the Basque models, that is roughly 180 to 365 years. If well-tuned code ran ten times faster, it is still 17 to 35 years. A cluster of a thousand data-centre GPUs would do it in a day or two. Same recipe, same code. The only difference is the compute.

SECTION TWELVE · FINE-TUNING

Fine-tuning a large model on your own data.

If the goal is a model that does a job for your organisation, training from scratch is rarely the route. Someone else has already paid for the expensive part. Gemma 4, released by Google under the Apache 2.0 licence, comes in four sizes, and you can take any of them and fine-tune it on your own documents, emails, tickets or reports. That is stage one of section ten, on a far stronger starting point, and it is well within reach of one desk-side machine.

The technique that makes it fit is LoRA. Instead of updating all of a model's weights, you freeze them and train a small set of extra weights alongside, typically well under one per cent of the total. QLoRA goes further and stores the frozen weights in 4-bit form, cutting their memory by about three-quarters. The result is an adapter file of a few hundred megabytes that changes how the base model behaves.

What fits on one DGX Spark, which has 128GB of memory:

ModelWhat it isFrozen weights, 4-bitFrozen weights, 16-bitFits on
Gemma 4 E2B2.3 billion effective parametersabout 3GBabout 10GBOne Spark, easily
Gemma 4 26B A4BMixture of experts: 25.2 billion parameters, 3.8 billion used per tokenabout 13GBabout 50GBOne Spark
Gemma 4 31BDense, 31 billion parametersabout 16GBabout 62GBOne Spark
Gemma 4 31B, every weight trainedFull fine-tune with the AdamW optimisernot applicableabout 500GB in totalFour linked Sparks, at the limit
Bar chart showing the memory needed to fine-tune Gemma 4 models: all LoRA and QLoRA cases fit well under one Spark's 128GB, and training every weight of the 31B model needs about 500GB, close to the 512GB of four linked Sparks. Gemma 4 E2B Gemma 4 26B A4B Gemma 4 31B Gemma 4 31B, every weight trained One Spark, 128GB Four linked Sparks, 512GB 3GB, 4-bit 10GB, 16-bit 13GB, 4-bit 50GB, 16-bit 16GB, 4-bit 62GB, 16-bit about 500GB, every weight 0 128 256 384 512 Memory for the frozen weights (GB)
Bar chart showing the memory needed to fine-tune Gemma 4 models: all LoRA and QLoRA cases fit well under one Spark's 128GB, and training every weight of the 31B model needs about 500GB, close to the 512GB of four linked Sparks. GEMMA 4 E2B 3GB, 4-bit 10GB, 16-bit GEMMA 4 26B A4B 13GB, 4-bit 50GB, 16-bit GEMMA 4 31B 16GB, 4-bit 62GB, 16-bit GEMMA 4 31B, EVERY WEIGHT TRAINED about 500GB, every weight One Spark, 128GB Four linked Sparks, 512GB 0 128 256 384 512 Memory for the frozen weights (GB)
Figure 4. Memory for the frozen weights when fine-tuning Gemma 4 models, against one DGX Spark and four linked Sparks. Training needs some headroom on top.

Those figures are the weights alone. Leave headroom for the working memory training needs, which grows with the length of the examples. NVIDIA's own guidance is that one 128GB Spark can fine-tune models of up to 70 billion parameters.

How long it takes. With LoRA, each training token costs about four times the model's active parameters in operations: the forward pass, plus working back through the frozen weights. For a company dataset of 10 million tokens, around 7.5 million words:

  • Gemma 4 31B: about 1.2 × 10¹⁸ operations. Roughly three days at the speed my Spark ran the Basque models, and well under a day with well-tuned code.
  • Gemma 4 26B A4B: only 3.8 billion parameters work on each token, so about 1.5 × 10¹⁷ operations, a matter of hours.

Two or three passes over the data multiply these. They are estimates from the same rule of thumb, not measurements, and 4-bit training runs somewhat slower than the arithmetic suggests. Measure your own throughput in the first few minutes and plan from that, as in section seven.

Linking Sparks. Up to four DGX Sparks can be linked through their 200Gb/s ConnectX-7 ports. Linking mainly buys memory: two give 256GB and four give 512GB, enough to train every weight of a 31-billion-parameter model, or to fine-tune on very long examples. It also shares out the work, but the link is far slower than the connections inside a data-centre server, so expect less than double the speed from two. For most fine-tuning, one Spark with LoRA is enough.

What fine-tuning is for. Fine-tuning changes how a model behaves: your house style, your document formats, your terminology, the shape of a good answer to your kind of question. It is a poor way to teach a model facts, especially facts that change. For questions like “what does our policy say”, retrieval, where the model looks up the relevant document at the moment it answers, is usually the better tool, and the two combine well. And for organisations that cannot send their data to a cloud provider, every step here, the data, the training and the finished model, stays inside the building.

SECTION THIRTEEN · THIS WINTER

What I am building this winter.

The most useful small model is often not a language model at all. It is a small model trained on one narrow job, with data only you have.

Mine is planned for this winter, for the boat. Over a season, the instruments log wind, boat speed and heading. A small model trained on those logs can learn how she actually sails, which is rarely what the published polars say, and give a live target speed at every wind angle. The second stage corrects the downloaded wind forecast for the waters I sail, by learning where it is reliably wrong. Both will run offline on a Raspberry Pi 5 with an AI accelerator board, with no signal needed at sea. I will write it up when it works.

SECTION FOURTEEN · COMMON QUESTIONS

Common questions.

Can anyone build their own LLM?

Yes. The method, the code and the data are public. A small model can be trained from scratch on one modern GPU in under a day, or over several nights on a home PC. What limits size and quality is compute and data.

Is building an AI model a secret or poorly understood process?

No. Every step, from the tokenizer to the training loop to chat tuning, is published and runs on open-source tools. What is still being researched is interpretability: explaining what a trained model's individual weights do.

How do you turn a language model into a chatbot?

Train the base model further on example conversations in a fixed format (supervised fine-tuning), then on pairs of better and worse answers (preference tuning), and wrap generation in a loop that keeps the conversation history.

Can you train a language model on a home PC?

Yes, at small sizes. Save checkpoints regularly and a run can be split across nights, starting when the PC is idle and resuming from the last checkpoint the next night. The GPU or CPU does the work, so idle processor time is what counts.

How much does it cost to train a language model?

It depends on size. GPT-2 small (124 million parameters) has been reproduced for about $20 of cloud GPU time. The largest Llama 3 model used 3.8 × 10²⁵ operations on Meta's own cluster, roughly a hundred million times the compute of a 110-million-parameter model trained for a day.

Should a business train its own model from scratch?

Rarely. Fine-tuning an existing open-weights model on your own data is faster and cheaper, and usually gives a better result for a specific task.

Can you fine-tune a large model on a DGX Spark?

Yes. With LoRA or QLoRA, a single 128GB DGX Spark can fine-tune open models such as Gemma 4 31B or Gemma 4 26B A4B on an organisation's own data, typically in hours to a few days for a dataset of around ten million tokens. Up to four Sparks can be linked for more memory.

CLOSING

How we work.

We build things to understand them, and we write up what we find. If you want to talk through where a model of your own would or would not help, the first conversation is a thirty-minute call, no pitch, no deck. We talk through where you are.