Skip to content
Part 4: On-Device iOS Deployment

Part 4: On-Device iOS Deployment

Running a fine-tuned model directly on an iPhone means zero latency to a server, offline capability, and complete privacy. The trade-off: you’re limited to small models (1–3B parameters) and the conversion pipeline has rough edges. This chapter covers both major paths: CoreML (Apple’s native ML framework) and MLX (Apple’s research ML framework for Apple Silicon).

The Landscape

Path Runtime Platform Maturity Best For
CoreML Core ML / ANE iOS, macOS, visionOS Production-ready Shipping apps to the App Store
MLX MLX Swift macOS, iOS (limited) Experimental Prototyping, research, macOS apps
llama.cpp GGML/Metal iOS, macOS Mature Maximum flexibility, C++ integration

For QuizMe on iOS, CoreML is the primary path. We’ll cover MLX as a secondary option useful for macOS prototyping and as the ecosystem matures.


Model Selection for On-Device

Not every model fits on a phone. Practical constraints:

Device RAM Realistic Model Size
iPhone 15 Pro 8 GB (shared) ≤ 3B quantized
iPhone 16 Pro 8 GB (shared) ≤ 3B quantized
iPhone 16 Pro Max 8 GB (shared) ≤ 3B quantized
iPad Pro M4 16 GB (shared) ≤ 7B quantized
MacBook (M-series) 16–128 GB Up to 70B quantized

Recommended base models for on-device QuizMe:

  1. LLaMA 3.2 1B / 3B: Meta designed these specifically for on-device. 1B is snappy, 3B is more capable.
  2. Phi-3.5 Mini (3.8B): Microsoft’s small model, strong reasoning for its size.
  3. SmolLM2 1.7B: Hugging Face’s purpose-built small model.
Strategy: Fine-tune LLaMA 3.2 3B for on-device (best quality/size ratio), and keep a larger model accessible via API as a fallback. The iOS app can try local first and fall back to the server for complex cases.

Path 1: CoreML Conversion

The Pipeline

Fine-tuned model (PyTorch) → Export to fp16 → Convert to CoreML → Quantize (4-bit) → Xcode integration

Step 1: Fine-Tune the Small Model

Repeat the Part 3 exercise, but with LLaMA 3.2 3B:

# Fine-tune with MLX LoRA (same approach as Part 3, smaller model)
mlx_lm.lora \
    --model mlx-community/Llama-3.2-3B-Instruct-4bit \
    --data data/quizme \
    --train \
    --batch-size 2 \
    --iters 1000 \
    --learning-rate 1e-5 \
    --lora-layers 16 \
    --adapter-path adapters/quizme-3b

# Fuse adapters into standalone model
mlx_lm.fuse \
    --model mlx-community/Llama-3.2-3B-Instruct-4bit \
    --adapter-path adapters/quizme-3b \
    --save-path models/quizme-3b-fused

Step 2: Convert to CoreML

Apple provides coremltools for conversion. The process goes through an intermediate format (either via exporters from Hugging Face or a direct torch trace).

uv pip install coremltools transformers torch
# convert_to_coreml.py
"""
Convert a merged Hugging Face model to CoreML format.

Note: As of early 2025, Apple's `ml-stable-diffusion` and
`swift-transformers` repos provide the most reliable conversion
paths. The landscape evolves quickly, so check Apple's docs.
"""
import coremltools as ct
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
import numpy as np

model_path = "outputs/quizme-3b/merged"

# Load the merged model
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
    model_path,
    torch_dtype=torch.float16,
    device_map="cpu",
)
model.eval()

# Trace the model
# CoreML needs a traced or scripted model
example_input = tokenizer(
    "Hello, world!",
    return_tensors="pt",
    max_length=128,
    padding="max_length",
)

traced_model = torch.jit.trace(
    model,
    example_kwarg_inputs={
        "input_ids": example_input["input_ids"],
        "attention_mask": example_input["attention_mask"],
    },
)

# Convert to CoreML
mlmodel = ct.convert(
    traced_model,
    inputs=[
        ct.TensorType(name="input_ids", shape=(1, 128), dtype=np.int32),
        ct.TensorType(name="attention_mask", shape=(1, 128), dtype=np.int32),
    ],
    minimum_deployment_target=ct.target.iOS18,
    compute_precision=ct.precision.FLOAT16,
)

# Quantize to 4-bit for on-device
mlmodel_quantized = ct.compression_utils.affine_quantize_weights(
    mlmodel,
    mode="linear_symmetric",
    dtype=ct.compression_utils.CompressionType.QUANTIZATION_4BIT,
)

mlmodel_quantized.save("outputs/quizme-3b/QuizMeGrader.mlpackage")
print("CoreML model saved!")

Reality check: Direct coremltools conversion of LLMs is fragile and often fails on custom architectures. The more reliable path as of 2025 is Apple’s exporters library (part of swift-transformers) or mlx-lm for MLX format. The code above illustrates the concept; you may need to use Apple’s specific export tooling depending on the model architecture.

Check these repos for the latest working pipelines:

Step 3: Integrate in Xcode (Swift)

Once you have the .mlpackage, add it to your Xcode project:

// QuizMeGrader.swift
import CoreML
import NaturalLanguage

class QuizMeGrader {
    private let model: QuizMeGrader_ML  // Auto-generated class from .mlpackage
    private let tokenizer: AutoTokenizer // From swift-transformers

    init() throws {
        let config = MLModelConfiguration()
        config.computeUnits = .all  // Use ANE + GPU + CPU
        self.model = try QuizMeGrader_ML(configuration: config)
        self.tokenizer = try AutoTokenizer.from(pretrained: "meta-llama/Llama-3.2-3B-Instruct")
    }

    func gradeAnswer(
        question: String,
        correctAnswer: String,
        studentAnswer: String,
        subject: String
    ) async throws -> GradeResult {
        let prompt = formatPrompt(
            question: question,
            correctAnswer: correctAnswer,
            studentAnswer: studentAnswer,
            subject: subject
        )

        let inputIds = tokenizer.encode(prompt)
        // Run autoregressive generation...
        // (Token-by-token generation loop using CoreML predictions)

        let output = tokenizer.decode(generatedIds)
        return try JSONDecoder().decode(GradeResult.self, from: output.data(using: .utf8)!)
    }

    private func formatPrompt(
        question: String,
        correctAnswer: String,
        studentAnswer: String,
        subject: String
    ) -> String {
        // Match the chat template used during fine-tuning
        return """
        <|start_header_id|>system<|end_header_id|>

        You are QuizMe, an expert quiz evaluator...<|eot_id|>
        <|start_header_id|>user<|end_header_id|>

        \(jsonInput)<|eot_id|>
        <|start_header_id|>assistant<|end_header_id|>

        """
    }
}

struct GradeResult: Codable {
    let score: Int
    let isCorrect: Bool
    let feedback: String
    let explanation: String

    enum CodingKeys: String, CodingKey {
        case score
        case isCorrect = "is_correct"
        case feedback
        case explanation
    }
}

Alternative: swift-transformers

For a higher-level approach, the swift-transformers package handles tokenization and generation:

import Hub
import Transformers

let hub = HubApi()
let modelId = "your-username/quizme-3b-coreml"  // Hosted on HF Hub

let model = try await AutoModelForCausalLM.from(
    pretrained: modelId,
    hubApi: hub
)

let tokenizer = try await AutoTokenizer.from(
    pretrained: modelId,
    hubApi: hub
)

let output = try await model.generate(
    tokenizer.encode(prompt),
    maxNewTokens: 512,
    temperature: 0.3
)

let text = tokenizer.decode(output)

Path 2: MLX (Apple Silicon Native)

MLX is Apple’s NumPy-like framework optimized for Apple Silicon’s unified memory architecture. It’s more mature for macOS than iOS, but the ecosystem is moving fast.

Convert to MLX Format

uv pip install mlx-lm
# Convert merged model to MLX format with 4-bit quantization
mlx_lm.convert \
    --hf-path outputs/quizme-3b/merged \
    --mlx-path outputs/quizme-3b/mlx \
    -q  # Quantize to 4-bit

Test with MLX

mlx_lm.generate \
    --model outputs/quizme-3b/mlx \
    --prompt '<|start_header_id|>system<|end_header_id|>

You are QuizMe...<|eot_id|><|start_header_id|>user<|end_header_id|>

{"question": "What is photosynthesis?", ...}<|eot_id|><|start_header_id|>assistant<|end_header_id|>

' \
    --max-tokens 512 \
    --temp 0.3

MLX Swift for macOS/iOS

// Using mlx-swift for on-device inference
import MLX
import MLXLLM

let model = try await LLMModel.load(from: "path/to/mlx/model")
let tokenizer = try await Tokenizer.load(from: "path/to/mlx/model")

let result = try await model.generate(
    prompt: formattedPrompt,
    maxTokens: 512,
    temperature: 0.3
)
MLX vs CoreML for iOS: CoreML has better hardware integration (accesses the Neural Engine directly), wider device support, and is the production-blessed path. MLX is better for rapid prototyping on macOS and gives you more control. For QuizMe, start with CoreML for the iOS app. Use MLX for testing on your Mac.

Path 3: llama.cpp (Pragmatic Alternative)

If CoreML conversion proves too fragile (it often does), llama.cpp via the GGUF format is the battle-tested alternative:

# You already exported GGUF in Part 3
# The file at outputs/quizme-3b/gguf/ can be used directly

iOS integration via llama.cpp’s Swift bindings:

// Using llama.cpp Swift wrapper
import llama

let model = try LlamaModel(path: "QuizMe-3B-Q4_K_M.gguf")
let result = try model.complete(
    prompt: formattedPrompt,
    maxTokens: 512,
    temperature: 0.3
)

This is the most reliable path for getting a model running on iOS today. The trade-off: no ANE acceleration (GPU + CPU only via Metal), so slightly slower than a well-optimized CoreML model.


Performance Expectations

Rough benchmarks for on-device generation (iPhone 15 Pro):

Model Format Tokens/sec First Token RAM Usage
LLaMA 3.2 1B CoreML 4-bit ~30-40 t/s ~0.5s ~1 GB
LLaMA 3.2 3B CoreML 4-bit ~15-20 t/s ~1s ~2.5 GB
LLaMA 3.2 3B GGUF Q4_K_M ~10-15 t/s ~1.5s ~2.5 GB
Phi-3 Mini 3.8B CoreML 4-bit ~12-18 t/s ~1.2s ~3 GB

For QuizMe’s JSON grading output (~200 tokens), expect 5–15 seconds for a complete response on the 3B model. Acceptable for an educational app, but consider showing a streaming UI.


Architecture Decision: Hybrid Approach

For a production QuizMe setup, the pragmatic architecture:

┌─────────────────────────────────────┐
│            QuizMe iOS App           │
├─────────────────────────────────────┤
│  1. Try local model (3B, CoreML)    │
│     └─ Fast, offline, private       │
│                                     │
│  2. If complex / low confidence:    │
│     └─ Call remote API (larger model│
│        └─ Better quality, online    │
│                                     │
│  3. Fallback: cloud API             │
│     └─ Claude/GPT for edge cases    │
└─────────────────────────────────────┘

The local model handles 80% of grading. Complex or ambiguous cases route to the larger remote model. Cloud API is the last resort.

Confidence detection: After generation, parse the JSON and check if the model’s score is in an “uncertain” range (e.g., 4–6) or if the output is malformed. These are good triggers to escalate to the larger model.

Further Reading

References