Part 4: On-Device iOS Deployment
Running a fine-tuned model directly on an iPhone means zero latency to a server, offline capability, and complete privacy. The trade-off: you’re limited to small models (1–3B parameters) and the conversion pipeline has rough edges. This chapter covers both major paths: CoreML (Apple’s native ML framework) and MLX (Apple’s research ML framework for Apple Silicon).
The Landscape
| Path | Runtime | Platform | Maturity | Best For |
|---|---|---|---|---|
| CoreML | Core ML / ANE | iOS, macOS, visionOS | Production-ready | Shipping apps to the App Store |
| MLX | MLX Swift | macOS, iOS (limited) | Experimental | Prototyping, research, macOS apps |
| llama.cpp | GGML/Metal | iOS, macOS | Mature | Maximum flexibility, C++ integration |
For QuizMe on iOS, CoreML is the primary path. We’ll cover MLX as a secondary option useful for macOS prototyping and as the ecosystem matures.
Model Selection for On-Device
Not every model fits on a phone. Practical constraints:
| Device | RAM | Realistic Model Size |
|---|---|---|
| iPhone 15 Pro | 8 GB (shared) | ≤ 3B quantized |
| iPhone 16 Pro | 8 GB (shared) | ≤ 3B quantized |
| iPhone 16 Pro Max | 8 GB (shared) | ≤ 3B quantized |
| iPad Pro M4 | 16 GB (shared) | ≤ 7B quantized |
| MacBook (M-series) | 16–128 GB | Up to 70B quantized |
Recommended base models for on-device QuizMe:
- LLaMA 3.2 1B / 3B: Meta designed these specifically for on-device. 1B is snappy, 3B is more capable.
- Phi-3.5 Mini (3.8B): Microsoft’s small model, strong reasoning for its size.
- SmolLM2 1.7B: Hugging Face’s purpose-built small model.
Path 1: CoreML Conversion
The Pipeline
Fine-tuned model (PyTorch) → Export to fp16 → Convert to CoreML → Quantize (4-bit) → Xcode integrationStep 1: Fine-Tune the Small Model
Repeat the Part 3 exercise, but with LLaMA 3.2 3B:
# Fine-tune with MLX LoRA (same approach as Part 3, smaller model)
mlx_lm.lora \
--model mlx-community/Llama-3.2-3B-Instruct-4bit \
--data data/quizme \
--train \
--batch-size 2 \
--iters 1000 \
--learning-rate 1e-5 \
--lora-layers 16 \
--adapter-path adapters/quizme-3b
# Fuse adapters into standalone model
mlx_lm.fuse \
--model mlx-community/Llama-3.2-3B-Instruct-4bit \
--adapter-path adapters/quizme-3b \
--save-path models/quizme-3b-fusedStep 2: Convert to CoreML
Apple provides coremltools for conversion. The process goes through an intermediate format (either via exporters from Hugging Face or a direct torch trace).
uv pip install coremltools transformers torch# convert_to_coreml.py
"""
Convert a merged Hugging Face model to CoreML format.
Note: As of early 2025, Apple's `ml-stable-diffusion` and
`swift-transformers` repos provide the most reliable conversion
paths. The landscape evolves quickly, so check Apple's docs.
"""
import coremltools as ct
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
import numpy as np
model_path = "outputs/quizme-3b/merged"
# Load the merged model
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.float16,
device_map="cpu",
)
model.eval()
# Trace the model
# CoreML needs a traced or scripted model
example_input = tokenizer(
"Hello, world!",
return_tensors="pt",
max_length=128,
padding="max_length",
)
traced_model = torch.jit.trace(
model,
example_kwarg_inputs={
"input_ids": example_input["input_ids"],
"attention_mask": example_input["attention_mask"],
},
)
# Convert to CoreML
mlmodel = ct.convert(
traced_model,
inputs=[
ct.TensorType(name="input_ids", shape=(1, 128), dtype=np.int32),
ct.TensorType(name="attention_mask", shape=(1, 128), dtype=np.int32),
],
minimum_deployment_target=ct.target.iOS18,
compute_precision=ct.precision.FLOAT16,
)
# Quantize to 4-bit for on-device
mlmodel_quantized = ct.compression_utils.affine_quantize_weights(
mlmodel,
mode="linear_symmetric",
dtype=ct.compression_utils.CompressionType.QUANTIZATION_4BIT,
)
mlmodel_quantized.save("outputs/quizme-3b/QuizMeGrader.mlpackage")
print("CoreML model saved!")Reality check: Direct coremltools conversion of LLMs is fragile and often fails on custom architectures. The more reliable path as of 2025 is Apple’s exporters library (part of swift-transformers) or mlx-lm for MLX format. The code above illustrates the concept; you may need to use Apple’s specific export tooling depending on the model architecture.
Check these repos for the latest working pipelines:
Step 3: Integrate in Xcode (Swift)
Once you have the .mlpackage, add it to your Xcode project:
// QuizMeGrader.swift
import CoreML
import NaturalLanguage
class QuizMeGrader {
private let model: QuizMeGrader_ML // Auto-generated class from .mlpackage
private let tokenizer: AutoTokenizer // From swift-transformers
init() throws {
let config = MLModelConfiguration()
config.computeUnits = .all // Use ANE + GPU + CPU
self.model = try QuizMeGrader_ML(configuration: config)
self.tokenizer = try AutoTokenizer.from(pretrained: "meta-llama/Llama-3.2-3B-Instruct")
}
func gradeAnswer(
question: String,
correctAnswer: String,
studentAnswer: String,
subject: String
) async throws -> GradeResult {
let prompt = formatPrompt(
question: question,
correctAnswer: correctAnswer,
studentAnswer: studentAnswer,
subject: subject
)
let inputIds = tokenizer.encode(prompt)
// Run autoregressive generation...
// (Token-by-token generation loop using CoreML predictions)
let output = tokenizer.decode(generatedIds)
return try JSONDecoder().decode(GradeResult.self, from: output.data(using: .utf8)!)
}
private func formatPrompt(
question: String,
correctAnswer: String,
studentAnswer: String,
subject: String
) -> String {
// Match the chat template used during fine-tuning
return """
<|start_header_id|>system<|end_header_id|>
You are QuizMe, an expert quiz evaluator...<|eot_id|>
<|start_header_id|>user<|end_header_id|>
\(jsonInput)<|eot_id|>
<|start_header_id|>assistant<|end_header_id|>
"""
}
}
struct GradeResult: Codable {
let score: Int
let isCorrect: Bool
let feedback: String
let explanation: String
enum CodingKeys: String, CodingKey {
case score
case isCorrect = "is_correct"
case feedback
case explanation
}
}Alternative: swift-transformers
For a higher-level approach, the swift-transformers package handles tokenization and generation:
import Hub
import Transformers
let hub = HubApi()
let modelId = "your-username/quizme-3b-coreml" // Hosted on HF Hub
let model = try await AutoModelForCausalLM.from(
pretrained: modelId,
hubApi: hub
)
let tokenizer = try await AutoTokenizer.from(
pretrained: modelId,
hubApi: hub
)
let output = try await model.generate(
tokenizer.encode(prompt),
maxNewTokens: 512,
temperature: 0.3
)
let text = tokenizer.decode(output)Path 2: MLX (Apple Silicon Native)
MLX is Apple’s NumPy-like framework optimized for Apple Silicon’s unified memory architecture. It’s more mature for macOS than iOS, but the ecosystem is moving fast.
Convert to MLX Format
uv pip install mlx-lm# Convert merged model to MLX format with 4-bit quantization
mlx_lm.convert \
--hf-path outputs/quizme-3b/merged \
--mlx-path outputs/quizme-3b/mlx \
-q # Quantize to 4-bitTest with MLX
mlx_lm.generate \
--model outputs/quizme-3b/mlx \
--prompt '<|start_header_id|>system<|end_header_id|>
You are QuizMe...<|eot_id|><|start_header_id|>user<|end_header_id|>
{"question": "What is photosynthesis?", ...}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
' \
--max-tokens 512 \
--temp 0.3MLX Swift for macOS/iOS
// Using mlx-swift for on-device inference
import MLX
import MLXLLM
let model = try await LLMModel.load(from: "path/to/mlx/model")
let tokenizer = try await Tokenizer.load(from: "path/to/mlx/model")
let result = try await model.generate(
prompt: formattedPrompt,
maxTokens: 512,
temperature: 0.3
)Path 3: llama.cpp (Pragmatic Alternative)
If CoreML conversion proves too fragile (it often does), llama.cpp via the GGUF format is the battle-tested alternative:
# You already exported GGUF in Part 3
# The file at outputs/quizme-3b/gguf/ can be used directlyiOS integration via llama.cpp’s Swift bindings:
// Using llama.cpp Swift wrapper
import llama
let model = try LlamaModel(path: "QuizMe-3B-Q4_K_M.gguf")
let result = try model.complete(
prompt: formattedPrompt,
maxTokens: 512,
temperature: 0.3
)This is the most reliable path for getting a model running on iOS today. The trade-off: no ANE acceleration (GPU + CPU only via Metal), so slightly slower than a well-optimized CoreML model.
Performance Expectations
Rough benchmarks for on-device generation (iPhone 15 Pro):
| Model | Format | Tokens/sec | First Token | RAM Usage |
|---|---|---|---|---|
| LLaMA 3.2 1B | CoreML 4-bit | ~30-40 t/s | ~0.5s | ~1 GB |
| LLaMA 3.2 3B | CoreML 4-bit | ~15-20 t/s | ~1s | ~2.5 GB |
| LLaMA 3.2 3B | GGUF Q4_K_M | ~10-15 t/s | ~1.5s | ~2.5 GB |
| Phi-3 Mini 3.8B | CoreML 4-bit | ~12-18 t/s | ~1.2s | ~3 GB |
For QuizMe’s JSON grading output (~200 tokens), expect 5–15 seconds for a complete response on the 3B model. Acceptable for an educational app, but consider showing a streaming UI.
Architecture Decision: Hybrid Approach
For a production QuizMe setup, the pragmatic architecture:
┌─────────────────────────────────────┐
│ QuizMe iOS App │
├─────────────────────────────────────┤
│ 1. Try local model (3B, CoreML) │
│ └─ Fast, offline, private │
│ │
│ 2. If complex / low confidence: │
│ └─ Call remote API (larger model│
│ └─ Better quality, online │
│ │
│ 3. Fallback: cloud API │
│ └─ Claude/GPT for edge cases │
└─────────────────────────────────────┘The local model handles 80% of grading. Complex or ambiguous cases route to the larger remote model. Cloud API is the last resort.
Further Reading
References
- Apple CoreML Documentation: Official CoreML docs.
- MLX Documentation: Apple’s MLX framework.
- mlx-lm: MLX LLM utilities (convert, quantize, serve).
- swift-transformers: Hugging Face’s Swift library for tokenization and generation.
- llama.cpp: The cross-platform LLM inference engine.
- llama.swiftui: llama.cpp SwiftUI example.