Skip to content

Training Overview

We’ll go much deeper into fine-tuning in Part 2 and get fully hands-on in Part 3. This page gives you the big picture: what are the different stages of training an LLM, and which ones do you actually do?

The Three Stages

Building a useful LLM happens in three stages, each building on the previous one:

Pre-Training  →  Supervised Fine-Tuning (SFT)  →  Alignment (RLHF / DPO)
   (months)           (hours to days)               (hours to days)

Stage 1: Pre-Training

Train on massive internet-scale text corpora (trillions of tokens). The objective is deceptively simple: predict the next token.

Given the sequence “The capital of France is,” the model should predict “Paris” with high probability. The loss function measures how wrong the prediction is:

L=−∑t=1Tlog⁡P(xt∣x<t)\mathcal{L} = -\sum_{t=1}^{T} \log P(x_t | x_{<t})

This is cross-entropy loss: for each position tt in the training text, take the log of the probability the model assigns to the actual next token. Sum them up, negate (because we minimize loss, and log of a probability is negative). Lower is better.

By doing this across trillions of tokens from books, websites, code, and conversations, the model learns:

  • Grammar and syntax
  • Factual knowledge (to the extent it’s in the training data)
  • Reasoning patterns
  • Code structure
  • Multiple languages

You don’t do this. Pre-training costs millions of dollars, takes months on thousands of GPUs, and requires enormous datasets. Meta, Google, Mistral, and others do this and release the resulting models for you to build on.

Stage 2: Supervised Fine-Tuning (SFT)

Take a pre-trained model and train it further on a smaller, curated dataset of (input, desired output) pairs. Same objective (predict the next token), but the data distribution is different: instead of random internet text, you’re training on examples of the specific behavior you want.

For example, to teach a model to follow instructions:

Input:  "Summarize this article in 3 bullet points: [article text]"
Output: "- Point one\n- Point two\n- Point three"

The model learns to associate the instruction pattern with the expected output format. After SFT on thousands of such examples, it reliably follows instructions rather than just completing text.

This is the stage you’ll do in Parts 3-4. For QuizMe, you’ll fine-tune a model on (question + student answer, grading JSON) pairs, teaching it to produce structured evaluations.

Key differences from pre-training:

  • Dataset size: Hundreds to tens of thousands of examples, not trillions of tokens
  • Learning rate: Much lower (we want to adjust, not overwrite what the model already knows)
  • Epochs: Small number (1-5), sometimes even less than one full pass
  • Cost: Minutes to hours on a single GPU (or your MacBook Pro)

Stage 3: Alignment (RLHF / DPO)

The final stage aligns the model’s behavior with human preferences. Even after SFT, a model might follow instructions but produce outputs that are unhelpful, harmful, or subtly wrong.

Alignment uses pairs of responses where one is preferred over the other:

Prompt:   "How do I pick a lock?"
Response A: [Detailed lock-picking instructions]
Response B: "I'd be happy to help with locksmithing! Here's how to contact a locksmith..."
Preferred: B

RLHF (Reinforcement Learning from Human Feedback) trains a separate “reward model” on these preferences, then uses it to guide the LLM. DPO (Direct Preference Optimization) is a simpler alternative that skips the reward model and directly optimizes from preference pairs.

This is what turns a raw pre-trained model into a helpful, harmless assistant. We’ll touch on DPO in Part 5 as an advanced technique for improving your fine-tuned model.


Where Fine-Tuning Fits

A helpful way to think about it:

Stage What it does Who does it Cost
Pre-training Gives the model general knowledge and language ability Meta, Google, Mistral, etc. Millions of dollars
SFT Teaches specific tasks, formats, or behaviors You Minutes to hours
Alignment Makes responses helpful and safe Model providers, or you (advanced) Hours to days

When you download a model like “LLaMA 3.2 3B Instruct,” stages 1-3 are already done. When you fine-tune it for QuizMe, you’re doing additional SFT on top of an already-instruction-tuned model. You’re not starting from scratch; you’re specializing.


Key Numbers to Internalize

These numbers help you develop intuition for model sizes and what fits where:

Model Parameters Layers d_model Heads Context Length
LLaMA 3.2 1B 1.24B 16 2048 32 128K
LLaMA 3.2 3B 3.21B 28 3072 24 128K
Mistral 7B 7.24B 32 4096 32 32K
LLaMA 3.1 8B 8.03B 32 4096 32 128K
Qwen 2.5 14B 14.8B 48 5120 40 128K
LLaMA 3.1 70B 70.6B 80 8192 64 128K

Memory Rules of Thumb

In fp16 (2 bytes per parameter):

Model size Weight memory Fits on MacBook Pro (16 GB)?
1B ~2 GB Yes, with room to spare
3B ~6 GB Yes, comfortably
7-8B ~14-16 GB Tight; quantization needed
14B ~28 GB Only with 4-bit quantization (~7 GB)
70B ~140 GB Only with aggressive quantization on 96+ GB machines

With 4-bit quantization (which we’ll use), memory drops ~4x:

Model size 4-bit memory Training overhead Total for fine-tuning
1B ~0.5 GB ~1-2 GB ~2-3 GB
3B ~1.5 GB ~2-3 GB ~4-6 GB
8B ~4 GB ~3-5 GB ~7-9 GB

This is why quantization matters, and why a 3B model is the sweet spot for fine-tuning on a MacBook Pro. We’ll cover quantization properly in Part 2.


Further Reading

References