Skip to content
Part 3: Hands-On Fine-Tuning

Part 3: Hands-On Fine-Tuning

Time to get your hands dirty. We’ll do two exercises:

  1. Generic: Fine-tune LLaMA 3.2 3B to follow a specific output format (structured movie reviews), which teaches the mechanics
  2. QuizMe: Fine-tune a model to grade quiz answers and explain correct answers, your real use case

Both use LoRA fine-tuning via MLX on your MacBook Pro.

Environment Setup

Create a Python environment

# Using uv (recommended)
mkdir ~/llm-fine-tuning && cd ~/llm-fine-tuning
uv init
uv venv --python 3.11
source .venv/bin/activate

Install MLX and dependencies

uv pip install "mlx-lm[train]"
uv pip install datasets wandb
mlx-lm is Apple’s framework for running and fine-tuning LLMs on Apple Silicon. It uses the Metal GPU directly through unified memory, so there’s no separate “GPU memory” to worry about. Your full system RAM is available to both the CPU and GPU.

Verify MLX access

import mlx.core as mx

print(f"MLX default device: {mx.default_device()}")
print(f"Metal available: {mx.metal.is_available()}")

Should show gpu as the default device and Metal available.

(Optional) Set up Weights & Biases

wandb login

This gives you training loss curves, gradient norms, and learning rate schedules in a web dashboard. Not required, but very helpful for debugging.


Exercise 1: Structured Movie Reviews

The goal: teach the model to output movie reviews in a strict JSON format with specific fields, given a free-form movie description.

Step 1: Prepare the Dataset

MLX expects JSONL files with one of three formats. For this exercise, we’ll use the completions format (prompt/completion pairs). Create a data/ directory with train.jsonl and valid.jsonl.

# prepare_movie_dataset.py
import json, random

examples = [
    {
        "prompt": "Write a structured review for the following movie.\n\nInception (2010) - A thief who steals corporate secrets through dream-sharing technology is given the inverse task of planting an idea into the mind of a C.E.O.",
        "completion": json.dumps({
            "title": "Inception",
            "year": 2010,
            "rating": 9,
            "genre": ["Sci-Fi", "Thriller", "Action"],
            "summary": "A visually stunning and intellectually ambitious film that layers dreams within dreams. Nolan delivers a heist movie wrapped in a philosophical puzzle about the nature of reality.",
            "strengths": ["Innovative concept", "Outstanding visual effects", "Complex but coherent plot", "Hans Zimmer's score"],
            "weaknesses": ["Emotional core could be stronger", "Some exposition-heavy dialogue"],
            "recommended_for": "Fans of cerebral sci-fi who enjoy films that reward multiple viewings"
        }, indent=2)
    },
    # ... add 50-200 more examples
]

# In production, generate diverse examples with Claude or GPT-4,
# then review and fix every one. We provide a fuller generator below.
Pro tip: Generate your training data with Claude or GPT-4. Write 10 examples by hand to set the standard, then use a stronger model to generate 100+ more following the same pattern. Review and fix every example. This is faster than writing all of them manually and gives you diversity.

For this exercise, we provide a ready-made dataset. Create generate_dataset.py:

# generate_dataset.py
"""
Generate a synthetic movie review dataset for fine-tuning.
In production, you'd curate these carefully or generate with a stronger model.
"""
import json, random, os

MOVIES = [
    ("The Shawshank Redemption", 1994, ["Drama"], 10,
     "Two imprisoned men bond over years, finding solace and eventual redemption through acts of common decency.",
     ["Powerful performances", "Masterful storytelling", "Emotional depth"],
     ["Slow pacing may not suit all viewers"]),
    ("Mad Max: Fury Road", 2015, ["Action", "Sci-Fi"], 9,
     "In a post-apocalyptic wasteland, a woman rebels against a tyrannical ruler with the help of a drifter.",
     ["Relentless practical action", "Strong visual storytelling", "Charlize Theron's performance"],
     ["Thin plot", "Limited dialogue"]),
    ("The Grand Budapest Hotel", 2014, ["Comedy", "Drama"], 8,
     "A writer encounters the owner of an aging hotel who tells of his early years as a lobby boy.",
     ["Unique visual style", "Witty dialogue", "Ensemble cast"],
     ["Style over substance for some", "Whimsical tone may not appeal to all"]),
    ("Parasite", 2019, ["Thriller", "Drama", "Comedy"], 10,
     "Greed and class discrimination threaten a symbiotic relationship between a wealthy family and a destitute clan.",
     ["Genre-defying storytelling", "Sharp social commentary", "Perfect tonal shifts"],
     ["Requires attention to subtitles"]),
    ("Blade Runner 2049", 2017, ["Sci-Fi", "Drama"], 8,
     "A young blade runner discovers a secret that could plunge society into chaos.",
     ["Stunning cinematography", "Thoughtful expansion of the original", "Villeneuve's direction"],
     ["Very long runtime", "Slow pacing"]),
    ("Everything Everywhere All at Once", 2022, ["Sci-Fi", "Comedy", "Action"], 9,
     "An aging Chinese immigrant is swept up in an insane adventure across the multiverse.",
     ["Wildly creative", "Emotional core", "Michelle Yeoh's performance", "Unique action sequences"],
     ["Overwhelming sensory experience", "Humor may not land for everyone"]),
    ("The Lighthouse", 2019, ["Horror", "Drama"], 7,
     "Two lighthouse keepers try to maintain sanity while stranded on a remote island.",
     ["Atmospheric cinematography", "Committed performances", "Unique vision"],
     ["Deliberately inaccessible", "Not for casual viewers", "Ambiguous to a fault"]),
    ("Arrival", 2016, ["Sci-Fi", "Drama"], 9,
     "A linguist works with the military to communicate with alien lifeforms after twelve mysterious spacecraft appear.",
     ["Intelligent sci-fi", "Amy Adams' performance", "Emotional twist ending", "Faithful adaptation"],
     ["Slow build may test patience", "Scientific liberties"]),
]

def make_example(movie):
    title, year, genres, rating, desc, strengths, weaknesses = movie
    rec_map = {10: "Essential viewing for any film enthusiast", 9: "Highly recommended for genre fans",
               8: "Worth watching, especially if you enjoy the genre", 7: "Recommended with reservations"}
    return {
        "prompt": f"Write a structured review for the following movie.\n\n{title} ({year}) - {desc}",
        "completion": json.dumps({
            "title": title, "year": year, "rating": rating,
            "genre": genres, "summary": desc,
            "strengths": strengths, "weaknesses": weaknesses,
            "recommended_for": rec_map.get(rating, "Genre enthusiasts")
        }, indent=2)
    }

all_examples = [make_example(m) for m in MOVIES]
# Duplicate with slight variations to pad the dataset for demonstration
# In real work, you'd have 100+ unique examples
all_examples = all_examples * 8  # 64 examples
random.shuffle(all_examples)

# Split: 80% train, 20% validation
split = int(len(all_examples) * 0.8)
train_set = all_examples[:split]
valid_set = all_examples[split:]

os.makedirs("data/movie-reviews", exist_ok=True)
for name, data in [("train", train_set), ("valid", valid_set)]:
    with open(f"data/movie-reviews/{name}.jsonl", "w") as f:
        for ex in data:
            f.write(json.dumps(ex) + "\n")

print(f"Generated {len(train_set)} train / {len(valid_set)} valid examples")

Step 2: Download and Convert the Model

MLX can use pre-converted models from Hugging Face (the mlx-community org hosts many), or you can convert one yourself.

# Option A: Use a pre-quantized model from mlx-community (fastest)
# The model downloads automatically on first use.

# Option B: Convert from HuggingFace yourself
mlx_lm.convert \
    --hf-path meta-llama/Llama-3.2-3B \
    --mlx-path models/llama-3.2-3b-mlx \
    --quantize \
    --q-bits 4 \
    --q-group-size 64
4-bit quantization cuts the model’s memory footprint roughly 4x. A 3B model in 4-bit uses about 1.5-2 GB of memory, leaving plenty of room for training on a 16 GB MacBook Pro.

Step 3: Fine-Tune with LoRA

CLI approach (simplest):

mlx_lm.lora \
    --model mlx-community/Llama-3.2-3B-4bit \
    --data data/movie-reviews \
    --train \
    --batch-size 4 \
    --iters 500 \
    --learning-rate 1e-5 \
    --lora-layers 16 \
    --adapter-path adapters/movie-reviews \
    --steps-per-report 10 \
    --steps-per-eval 50 \
    --val-batches 10

Python API (more control):

# train_movie_reviews.py
import mlx.core as mx
import mlx.optimizers as optim
from mlx_lm import load
from mlx_lm.tuner import train, TrainingArgs
from mlx_lm.tuner.datasets import load_dataset
from mlx_lm.tuner.utils import linear_to_lora_layers

# Load model + tokenizer
model, tokenizer = load("mlx-community/Llama-3.2-3B-4bit")

# Freeze base model, apply LoRA to the last 16 layers
model.freeze()
lora_params = {
    "rank": 16,           # LoRA rank (same concept as Part 2)
    "dropout": 0.05,
    "scale": 32.0,        # Equivalent to alpha in PyTorch LoRA
}
linear_to_lora_layers(model, num_layers=16, lora_parameters=lora_params)

# Count trainable parameters
trainable = sum(p.size for n, p in model.trainable_parameters().items())
total = sum(p.size for n, p in model.parameters().items())
print(f"Trainable: {trainable:,} / {total:,} ({100 * trainable / total:.2f}%)")

# Load dataset (expects train.jsonl + valid.jsonl in the directory)
train_set, valid_set, _ = load_dataset(
    "data/movie-reviews",
    tokenizer=tokenizer,
)

# Configure training
training_args = TrainingArgs(
    batch_size=4,
    iters=500,
    val_batches=10,
    steps_per_report=10,
    steps_per_eval=50,
    steps_per_save=100,
    adapter_file="adapters/movie-reviews/adapters.safetensors",
    max_seq_length=2048,
)

optimizer = optim.Adam(learning_rate=1e-5)

# Train
train(
    model=model,
    args=training_args,
    optimizer=optimizer,
    train_dataset=train_set,
    val_dataset=valid_set,
)

print("Training complete!")

Step 4: Test the Model

# test_movie_reviews.py
from mlx_lm import load, generate

# Load base model + LoRA adapters
model, tokenizer = load(
    "mlx-community/Llama-3.2-3B-4bit",
    adapter_path="adapters/movie-reviews",
)

prompt = """Write a structured review for the following movie.

Dune: Part Two (2024) - Paul Atreides unites with the Fremen to seek revenge against those who destroyed his family while trying to prevent a terrible future only he can foresee."""

response = generate(
    model, tokenizer,
    prompt=prompt,
    max_tokens=512,
    temp=0.7,
    top_p=0.9,
)

print(response)

Or via the CLI:

mlx_lm.generate \
    --model mlx-community/Llama-3.2-3B-4bit \
    --adapter-path adapters/movie-reviews \
    --prompt "Write a structured review for the following movie.

Dune: Part Two (2024) - Paul Atreides unites with the Fremen..." \
    --max-tokens 512 \
    --temp 0.7

Step 5: Save the Model

# Fuse LoRA adapters into the base model (creates a standalone model)
mlx_lm.fuse \
    --model mlx-community/Llama-3.2-3B-4bit \
    --adapter-path adapters/movie-reviews \
    --save-path models/movie-reviews-fused

# Optional: export to GGUF for Ollama / llama.cpp
mlx_lm.fuse \
    --model mlx-community/Llama-3.2-3B-4bit \
    --adapter-path adapters/movie-reviews \
    --save-path models/movie-reviews-gguf \
    --export-gguf

Exercise 2: QuizMe Answer Grading

Now let’s apply the same process to your actual use case: a model that evaluates student answers and explains the correct answer.

Define the Task

The model receives:

  • A quiz question
  • The correct answer
  • The student’s answer
  • The subject/topic

It outputs:

  • A score (0-10)
  • Whether the answer is correct (boolean)
  • A brief explanation of why the score was given
  • An explanation of the correct answer (for learning)

Step 1: Design the Output Schema

{
  "score": 7,
  "is_correct": true,
  "feedback": "Your answer captures the core concept but lacks precision. You correctly identified gravity as the cause but didn't mention the specific mechanism.",
  "explanation": "Tides are primarily caused by the gravitational pull of the Moon on Earth's oceans. The Moon's gravity creates two tidal bulges: one on the side of Earth nearest the Moon (direct gravitational pull) and one on the opposite side (due to inertia). The Sun also contributes, creating spring and neap tides depending on alignment."
}

Step 2: Build the Dataset

This is the most important step. We’ll use the chat format (JSONL with messages field) since we’re building on an instruct model.

# generate_quizme_dataset.py
"""
Generate training data for the QuizMe answer grading model.
Strategy:
  1. Define question-answer pairs across subjects
  2. For each, create student answers at different quality levels
  3. Ensure consistent grading calibration
"""
import json, os

# Calibration rubric:
# 10: Perfect, comprehensive answer
# 8-9: Correct with minor gaps
# 5-7: Partially correct, key concepts present but incomplete
# 3-4: Shows some understanding but significant errors
# 1-2: Mostly wrong, minimal understanding shown
# 0: Completely wrong or nonsensical

SYSTEM_PROMPT = "You are QuizMe, an expert quiz evaluator. Given a question, the correct answer, and a student's answer, provide a score (0-10), correctness assessment, specific feedback, and an educational explanation of the correct answer. Respond in JSON format."

dataset = []

# --- Science ---
dataset.extend([
    {
        "messages": [
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": json.dumps({
                "question": "What causes tides on Earth?",
                "correct_answer": "Tides are primarily caused by the gravitational pull of the Moon on Earth's oceans, with a secondary contribution from the Sun. The Moon's gravity creates tidal bulges on both the near and far sides of Earth.",
                "student_answer": "The moon pulls on the water.",
                "subject": "Earth Science"
            })},
            {"role": "assistant", "content": json.dumps({
                "score": 5,
                "is_correct": False,
                "feedback": "You've identified the core mechanism (the Moon's gravitational pull) but the answer is too vague to demonstrate real understanding. To improve, mention that the effect creates TWO tidal bulges (near and far side), and that the Sun also plays a role.",
                "explanation": "Tides are primarily caused by the Moon's gravitational attraction on Earth's oceans. This creates two tidal bulges: one facing the Moon (direct pull) and one on the opposite side (centrifugal effect of the Earth-Moon orbital system). The Sun contributes about 46% as much tidal force as the Moon. When Sun, Moon, and Earth align (new/full moon), we get especially strong 'spring tides.' When they're at right angles (quarter moons), we get weaker 'neap tides.'"
            }, indent=2)}
        ]
    },
    {
        "messages": [
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": json.dumps({
                "question": "What causes tides on Earth?",
                "correct_answer": "Tides are primarily caused by the gravitational pull of the Moon on Earth's oceans, with a secondary contribution from the Sun. The Moon's gravity creates tidal bulges on both the near and far sides of Earth.",
                "student_answer": "Tides happen because the Moon's gravity pulls on Earth's oceans, creating a bulge of water on the side closest to the Moon. There's also a bulge on the opposite side due to centrifugal force. The Sun has a smaller effect too.",
                "subject": "Earth Science"
            })},
            {"role": "assistant", "content": json.dumps({
                "score": 9,
                "is_correct": True,
                "feedback": "Excellent answer! You correctly identified both tidal bulges and mentioned the Sun's contribution. The only minor note: the far-side bulge is better explained by the differential in gravitational force across Earth's diameter rather than purely 'centrifugal force,' though your understanding is functionally correct.",
                "explanation": "Tides are primarily caused by the Moon's gravitational attraction on Earth's oceans. This creates two tidal bulges: one facing the Moon (direct pull) and one on the opposite side (centrifugal effect of the Earth-Moon orbital system). The Sun contributes about 46% as much tidal force as the Moon. When Sun, Moon, and Earth align (new/full moon), we get especially strong 'spring tides.' When they're at right angles (quarter moons), we get weaker 'neap tides.'"
            }, indent=2)}
        ]
    },
    {
        "messages": [
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": json.dumps({
                "question": "What causes tides on Earth?",
                "correct_answer": "Tides are primarily caused by the gravitational pull of the Moon on Earth's oceans, with a secondary contribution from the Sun. The Moon's gravity creates tidal bulges on both the near and far sides of Earth.",
                "student_answer": "Wind pushes the water around.",
                "subject": "Earth Science"
            })},
            {"role": "assistant", "content": json.dumps({
                "score": 0,
                "is_correct": False,
                "feedback": "This is incorrect. Wind causes waves, not tides. Tides are a gravitational phenomenon caused by the Moon and Sun, not by weather. The key distinction: waves are irregular surface motion from wind; tides are regular, predictable rises and falls of sea level on roughly a 12.5-hour cycle.",
                "explanation": "Tides are primarily caused by the Moon's gravitational attraction on Earth's oceans. This creates two tidal bulges: one facing the Moon (direct pull) and one on the opposite side (centrifugal effect of the Earth-Moon orbital system). The Sun contributes about 46% as much tidal force as the Moon. When Sun, Moon, and Earth align (new/full moon), we get especially strong 'spring tides.' When they're at right angles (quarter moons), we get weaker 'neap tides.'"
            }, indent=2)}
        ]
    },
])

# Add more question sets across subjects:
# - History, Math, Biology, Literature, Geography, Computer Science, etc.
# - Each question should have 3-5 student answers at different quality levels
# - Target: 200-500 total examples

os.makedirs("data/quizme", exist_ok=True)

# Split 80/20
split = max(1, int(len(dataset) * 0.8))
for name, data in [("train", dataset[:split]), ("valid", dataset[split:])]:
    with open(f"data/quizme/{name}.jsonl", "w") as f:
        for ex in data:
            f.write(json.dumps(ex) + "\n")

print(f"Generated {len(dataset)} total examples")
print("NOTE: Expand to 200+ examples across 20+ subjects for a production model.")
The dataset above is a skeleton. A production-quality dataset needs 200-500 examples minimum, across 15-20+ subjects, with 3-5 answer quality levels per question. Use Claude (via the API or Claude Code) to generate diverse examples, then review and calibrate every one. The grading must be internally consistent: a “7” in science should feel like a “7” in history.

Step 3: Train with Chat Format

CLI approach:

mlx_lm.lora \
    --model mlx-community/Llama-3.2-3B-Instruct-4bit \
    --data data/quizme \
    --train \
    --batch-size 2 \
    --iters 1000 \
    --learning-rate 1e-5 \
    --lora-layers 16 \
    --lora-rank 32 \
    --adapter-path adapters/quizme \
    --steps-per-report 10 \
    --steps-per-eval 100 \
    --val-batches 10

Python API:

# train_quizme.py
import mlx.core as mx
import mlx.optimizers as optim
from mlx_lm import load
from mlx_lm.tuner import train, TrainingArgs
from mlx_lm.tuner.datasets import load_dataset
from mlx_lm.tuner.utils import linear_to_lora_layers

# Load instruct model (we want to preserve instruction-following)
model, tokenizer = load("mlx-community/Llama-3.2-3B-Instruct-4bit")

# Apply LoRA
model.freeze()
lora_params = {
    "rank": 32,           # Higher rank for a more complex task
    "dropout": 0.05,
    "scale": 64.0,
}
linear_to_lora_layers(model, num_layers=16, lora_parameters=lora_params)

# Check trainable parameters
trainable = sum(p.size for n, p in model.trainable_parameters().items())
total = sum(p.size for n, p in model.parameters().items())
print(f"Trainable: {trainable:,} / {total:,} ({100 * trainable / total:.2f}%)")

# Load dataset (chat format with "messages" field)
train_set, valid_set, _ = load_dataset(
    "data/quizme",
    tokenizer=tokenizer,
)

# Configure training
training_args = TrainingArgs(
    batch_size=2,
    iters=1000,
    val_batches=10,
    steps_per_report=10,
    steps_per_eval=100,
    steps_per_save=200,
    adapter_file="adapters/quizme/adapters.safetensors",
    max_seq_length=2048,
)

optimizer = optim.Adam(learning_rate=1e-5)

# Train
train(
    model=model,
    args=training_args,
    optimizer=optimizer,
    train_dataset=train_set,
    val_dataset=valid_set,
)

# Fuse and save
print("Training complete. Fuse adapters with:")
print("  mlx_lm.fuse --model mlx-community/Llama-3.2-3B-Instruct-4bit \\")
print("    --adapter-path adapters/quizme --save-path models/quizme-fused")

Step 4: Test QuizMe

# test_quizme.py
from mlx_lm import load, generate
import json

# Load model with adapters
model, tokenizer = load(
    "mlx-community/Llama-3.2-3B-Instruct-4bit",
    adapter_path="adapters/quizme",
)

# Test with a new question
test_input = {
    "question": "Explain the difference between a stack and a queue.",
    "correct_answer": "A stack is a LIFO (Last-In, First-Out) data structure where elements are added and removed from the same end (top). A queue is a FIFO (First-In, First-Out) data structure where elements are added at the rear and removed from the front.",
    "student_answer": "A stack is like a pile of plates where you take from the top. A queue is like a line at the store where the first person gets served first. Stack is LIFO and queue is FIFO.",
    "subject": "Computer Science"
}

messages = [
    {"role": "system", "content": "You are QuizMe, an expert quiz evaluator. Given a question, the correct answer, and a student's answer, provide a score (0-10), correctness assessment, specific feedback, and an educational explanation of the correct answer. Respond in JSON format."},
    {"role": "user", "content": json.dumps(test_input)},
]

# Apply chat template
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

response = generate(
    model, tokenizer,
    prompt=prompt,
    max_tokens=512,
    temp=0.3,  # Low temperature for consistent grading
    top_p=0.9,
)

print(response)

# Validate JSON output
try:
    result = json.loads(response)
    print(f"\nScore: {result['score']}/10")
    print(f"Correct: {result['is_correct']}")
    print(f"Feedback: {result['feedback']}")
except json.JSONDecodeError:
    print("WARNING: Model did not produce valid JSON.")
    print("Dataset needs more examples or training needs more iterations.")

Step 5: Export for Deployment

# Fuse adapters into standalone model
mlx_lm.fuse \
    --model mlx-community/Llama-3.2-3B-Instruct-4bit \
    --adapter-path adapters/quizme \
    --save-path models/quizme-fused

# Export to GGUF for Ollama / llama.cpp (useful for iOS via llama.cpp)
mlx_lm.fuse \
    --model mlx-community/Llama-3.2-3B-Instruct-4bit \
    --adapter-path adapters/quizme \
    --save-path models/quizme-gguf \
    --export-gguf

# Optional: serve locally with an OpenAI-compatible API
mlx_lm.server \
    --model mlx-community/Llama-3.2-3B-Instruct-4bit \
    --adapter-path adapters/quizme \
    --port 8080

The local server exposes POST /v1/chat/completions, so you can test with curl or any OpenAI-compatible client before moving to on-device deployment in Part 4.


Troubleshooting

Common issues and fixes:

Out of memory: Reduce --batch-size to 1, reduce --max-seq-length, or use a smaller model. On a 16 GB MacBook Pro, a 3B model in 4-bit with LoRA training should use around 6-8 GB.

Loss not decreasing: Learning rate might be too low or too high. Start with 1e-5 for quantized models. Inspect your data for inconsistencies.

Loss drops then spikes: Overfitting. Reduce --iters, add more diverse training data, or increase LoRA dropout.

Model outputs garbage after fine-tuning: Chat template mismatch. Verify you’re using the same template for training and inference. Use tokenizer.apply_chat_template() consistently.

JSON output malformed: The model needs more examples with consistent JSON formatting. Add 50+ more training examples with valid JSON. Lower temperature during inference (0.1-0.3).

Slow training: Make sure you’re running on Apple Silicon (not under Rosetta). Check that mx.metal.is_available() returns True. Close other memory-heavy applications.


What’s Next

You now have a fine-tuned model that can grade quiz answers. In Part 4, we’ll convert it for on-device inference on iOS. In Part 5, we’ll build evaluation pipelines to systematically measure how well it’s actually performing.

Exercise Files

All scripts from this chapter:

  • generate_dataset.py – Synthetic movie review dataset generator
  • train_movie_reviews.py – Exercise 1 training script (Python API)
  • generate_quizme_dataset.py – QuizMe dataset generator (skeleton)
  • train_quizme.py – Exercise 2 training script
  • test_quizme.py – QuizMe inference test