Part 2: Models & Fine-Tuning Concepts
Now that you know how a transformer works, let’s talk about the landscape of models available, what it means to fine-tune one, and the techniques that make it practical on consumer hardware.
The Model Landscape
Open-Weight Model Families
The fine-tuning ecosystem revolves around a handful of model families. Here’s what matters:
Meta LLaMA 3.x: The de facto standard for fine-tuning. Strong base models at 1B, 3B, 8B, and 70B. Huge community, excellent tooling support. The 8B model hits a sweet spot for quality vs. hardware requirements. LLaMA 3.2 added small models (1B, 3B) specifically designed for on-device and edge deployment, directly relevant to Part 4.
Mistral / Mixtral: Mistral 7B punches above its weight through sliding window attention and GQA. Mixtral is a Mixture-of-Experts (MoE) architecture: 8 expert FFN blocks per layer, but only 2 are active per token, giving you 47B total parameters but only ~13B active during inference.
Qwen 2.5: Alibaba’s series. Notably strong multilingual performance (CJK languages especially). Available at 0.5B through 72B. Good option if your use case involves non-English text.
Microsoft Phi: Small models (1.3B to 14B) trained on high-quality “textbook” data. Phi-3/4 Mini at 3.8B is remarkably capable for its size. Interesting for on-device scenarios.
Google Gemma 2: 2B, 9B, 27B. Solid models, less community fine-tuning tooling than LLaMA.
SmolLM2: Hugging Face’s small model series (135M to 1.7B). Purpose-built for on-device use.
Base vs. Instruct Models
Every model family publishes two variants:
-
Base model: Pre-trained on next-token prediction only. It’s a text completion engine. Feed it “The capital of France is” and it’ll say “Paris.” Feed it “What is the capital of France?” and it might continue with another question, because questions in its training data are often followed by more questions.
-
Instruct model: Base model + supervised fine-tuning (SFT) on instruction/response pairs + alignment (RLHF or DPO). This is what makes a model follow instructions rather than just complete text.
For fine-tuning, you usually start from the base model if you’re training on a specific task format, or from the instruct model if you want to preserve general instruction-following and add domain knowledge on top.
What Fine-Tuning Actually Means
Fine-tuning is continued training on a smaller, task-specific dataset. The model’s weights are updated via the same gradient descent process used in pre-training, but:
- The dataset is much smaller (hundreds to tens of thousands of examples, not trillions of tokens)
- The learning rate is much lower (we want to adjust, not overwrite)
- The number of epochs is small (1–5 typically)
What Fine-Tuning Is Good For
- Teaching a specific output format (JSON, structured grading rubrics, XML tags)
- Instilling domain knowledge (medical terminology, legal reasoning patterns)
- Adjusting tone and style (formal reports, casual conversation, a specific persona)
- Improving performance on a narrow task (classification, extraction, grading)
- Reducing token usage, since a fine-tuned model often needs less prompting because the behavior is baked in
What Fine-Tuning Is NOT Good For
- Adding factual knowledge: the model has limited capacity to memorize new facts through fine-tuning. Use RAG for knowledge retrieval.
- Fixing fundamental reasoning limitations: if a 7B model can’t solve a class of problems with good prompting, fine-tuning probably won’t fix it. Try a larger model.
- General-purpose improvement: you’ll typically make the model better at your task and slightly worse at others (catastrophic forgetting).
The Decision Framework
Can you solve it with prompt engineering alone?
├─ Yes → Don't fine-tune. Ship the prompt.
└─ No → Does the model need to know things it doesn't?
├─ Yes → RAG (retrieval-augmented generation)
└─ No → Does the model need to behave differently?
├─ Yes → Fine-tune
└─ No → Try a larger/better modelQuantization
Quantization reduces the numerical precision of model weights, trading a small amount of quality for dramatic memory savings.
Precision Formats
| Format | Bits/Param | Memory for 7B | Notes |
|---|---|---|---|
| fp32 | 32 | ~28 GB | Full precision, rarely used for inference |
| fp16 / bf16 | 16 | ~14 GB | Standard training precision |
| int8 | 8 | ~7 GB | Good quality, ~2× compression |
| int4 (NF4) | 4 | ~3.5 GB | Used by QLoRA, ~4× compression |
| GGUF Q4_K_M | ~4.5 | ~4 GB | llama.cpp format, optimized for CPU+GPU |
Quantization Methods
bitsandbytes (bnb): The most common for fine-tuning on NVIDIA hardware. Provides 4-bit NormalFloat (NF4) quantization, which is theoretically optimal for normally distributed weights. This is what QLoRA uses. NVIDIA-only; on Apple Silicon, MLX handles quantized fine-tuning natively instead.
GPTQ: Post-training quantization. Quantizes layer-by-layer using a calibration dataset. Produces static quantized weights. Good for inference, not used during training.
AWQ (Activation-Aware Weight Quantization): Similar to GPTQ but preserves the weights that matter most for activations. Slightly better quality than GPTQ at the same bit width.
GGUF: The llama.cpp format. Supports mixed quantization (different layers get different precision). The Q4_K_M variant is the sweet spot for quality vs. size. Designed for CPU inference with optional GPU offloading.
Parameter-Efficient Fine-Tuning (PEFT)
Full fine-tuning updates every weight in the model. For a 7B model in fp16, that means storing gradients and optimizer states for 7 billion parameters, easily 60+ GB of memory. Not happening on a MacBook Pro, even with unified memory.
PEFT methods freeze most weights and only train a small number of additional parameters. The big ones:
LoRA (Low-Rank Adaptation)
The core insight: the weight updates during fine-tuning occupy a low-rank subspace. Instead of updating the full weight matrix , LoRA decomposes the update into two small matrices:
where and , with rank .
For a 4096 × 4096 weight matrix:
- Full fine-tuning: 16.7M parameters to update
- LoRA with : parameters, 128x fewer
Key hyperparameters:
- Rank (): Typically 8-64. Higher rank = more expressive but more parameters. Start with 16.
- Alpha (): Scaling factor. The actual update is . Common to set .
- Target modules: Which weight matrices to apply LoRA to. At minimum: Q and V projections in attention. For better results: Q, K, V, O (all attention) + gate, up, down (FFN projections).
Initialization: is initialized with random Gaussian values, is initialized to zero. This means the LoRA update starts at zero and the model behaves identically to the frozen base model at the start of training.
QLoRA (Quantized LoRA)
QLoRA combines LoRA with 4-bit quantization:
- Quantize the base model to 4-bit NF4 (using bitsandbytes)
- Freeze the quantized weights
- Add LoRA adapters in fp16/bf16 on top
- Train only the LoRA adapters
Memory footprint for a 7B model:
- Base model (4-bit): ~3.5 GB
- LoRA adapters (bf16): ~50–200 MB
- Gradients + optimizer states: ~1–2 GB
- Total: ~5-6 GB (fits easily in 16 GB of unified memory)
This is the technique we’ll use in Part 3. The quality loss from 4-bit quantization during training is surprisingly small, especially with NF4.
Other PEFT Methods
- DoRA (Weight-Decomposed LoRA): Decomposes weight updates into magnitude and direction. Slightly better than LoRA on benchmarks, compatible with the same tooling.
- IA3: Scales activations with three learned vectors. Fewer parameters than LoRA but less flexible.
- Prompt Tuning / Prefix Tuning: Prepend learnable “virtual tokens” to the input. Simple but less effective than LoRA for most tasks.
Just use QLoRA. It’s the best trade-off of quality, memory efficiency, and tooling maturity.
Dataset Formats
Fine-tuning is only as good as your data. Here are the standard formats:
Alpaca Format
Simple instruction/input/output triples. Good for single-turn tasks:
{
"instruction": "Grade the following answer on a scale of 0-10.",
"input": "Question: What causes tides?\nAnswer: The moon's gravity pulls on the ocean.",
"output": "Score: 7/10\nThe answer correctly identifies the moon's gravity as the primary cause. It would be stronger if it mentioned the sun's gravitational contribution and the concept of tidal bulges on opposite sides of Earth."
}ShareGPT / Conversational Format
Multi-turn conversations. Used for chat-style fine-tuning:
{
"conversations": [
{ "role": "system", "content": "You are an expert quiz evaluator." },
{ "role": "user", "content": "Grade this answer: ..." },
{ "role": "assistant", "content": "Score: 7/10\n..." }
]
}Chat Templates
Each model family has its own chat template: the special tokens and formatting that delineate system, user, and assistant messages. For LLaMA 3:
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
You are a helpful assistant.<|eot_id|><|start_header_id|>user<|end_header_id|>
Hello!<|eot_id|><|start_header_id|>assistant<|end_header_id|>
Hi there!<|eot_id|>Getting the chat template wrong will ruin your fine-tune. This is one of the most common mistakes. MLX’s load_dataset and tokenizer.apply_chat_template() handle this automatically when you use the chat JSONL format.
Dataset Quality > Quantity
Principles for dataset construction:
- Diversity: Cover the range of inputs the model will see in production. Edge cases, different phrasings, varying difficulty levels.
- Consistency: If your output format is “Score: X/10” followed by an explanation, every example should follow that format exactly.
- Correctness: Every example should be a gold-standard response you’d be happy to ship.
- Realistic difficulty: Include easy, medium, and hard examples in realistic proportions.
For the QuizMe use case, this means: varied subjects, different answer quality levels (perfect, partial, wrong, nonsensical), consistent grading format, calibrated scores.
Training Hyperparameters Cheat Sheet
These are the knobs you’ll turn in Part 3. Defaults here are sensible starting points:
| Parameter | Typical Range | Notes |
|---|---|---|
| Learning rate | 1e-4 to 2e-4 | Lower than pre-training. 2e-4 is standard for QLoRA. |
| Batch size | 4–16 (effective) | Use gradient accumulation to simulate larger batches. |
| Epochs | 1–5 | Monitor eval loss; stop if it starts rising. |
| Max sequence length | 512–2048 | Depends on your data. Longer = more VRAM. |
| Warmup steps | 5–10% of total | Linear warmup then cosine/linear decay. |
| Weight decay | 0.01–0.1 | Regularization. Start with 0.01. |
| LoRA rank () | 8–64 | 16 is the default starting point. |
| LoRA alpha () | Scaling factor. | |
| LoRA dropout | 0.05–0.1 | Regularization on LoRA layers. |
Merging and Export
After training, you have:
- The frozen, quantized base model
- A small set of LoRA adapter weights (~50–200 MB)
You can either:
- Serve with adapters loaded: Load base model + adapter at inference. Allows hot-swapping adapters for different tasks on the same base model.
- Merge adapters into the base model: Dequantize base weights to fp16, add the LoRA update, save as a single model. Simpler for deployment, required for format conversion (CoreML, GGUF).
Merging formula:
Further Reading
References
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., 2021): The original LoRA paper.
- QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al., 2023): Introduced NF4 quantization + double quantization + paged optimizers.
- A Visual Guide to Quantization: Maarten Grootendorst’s excellent visual explainer.
- Hugging Face PEFT Documentation: The library you’ll use.
- MLX LM Documentation: The fine-tuning framework for Part 3 (Apple Silicon native).