Part 5: LLM Ops & Evaluation
You’ve fine-tuned a model. How do you know it’s actually good? How do you monitor it in production? How do you systematically improve it over time? This chapter covers the concepts and techniques for LLM evaluation, observability, and continuous improvement.
The Evaluation Problem
Traditional software has clear pass/fail tests. LLMs don’t. A grading model might give a “7/10” where a human would give “6/10”: is that wrong? What about “8/10”? You need frameworks for measuring quality that go beyond exact match.
Evaluation Dimensions
For a model like QuizMe, evaluation isn’t a single number. You care about several dimensions independently:
| Dimension | What it measures | How to check |
|---|---|---|
| Score accuracy | Does the model’s score match human-calibrated scores? | Compare against gold-standard dataset. ±1 is acceptable, ±3 is a problem. |
| Format compliance | Does it output valid JSON with all required fields? | Automated parsing check on every output. |
| Feedback quality | Is the feedback specific, actionable, and factually correct? | LLM-as-judge or human review. |
| Explanation quality | Is the explanation accurate and educational? | LLM-as-judge or human review. |
| Consistency | Same input, same score across runs? | Run the same examples multiple times. Low temperature helps. |
| Calibration | Is a “7” in science comparable to a “7” in history? | Cross-subject analysis on the eval dataset. |
A model can score well on format compliance and consistency but poorly on feedback quality. Evaluating along each dimension separately tells you where to improve, not just whether to improve.
Observability Platforms
Once your model is running (whether on-device or via API), you need visibility into what it’s doing. LLM observability platforms record every call and let you analyze quality, cost, and performance over time.
Langfuse
Langfuse is the leading open-source LLM observability platform. It can be self-hosted or used as a cloud service.
Core concepts:
- Traces: A record of every LLM call, including input, output, latency, token usage, and cost. Think of it like structured logging, but purpose-built for LLM interactions.
- Scores: Quality ratings attached to traces. These can come from three sources: automated checks (did the JSON parse?), LLM-as-judge ratings, or human feedback.
- Datasets: Curated sets of inputs with expected outputs. You run your model against these to measure quality across versions.
- Experiments: Side-by-side comparisons of model versions on the same dataset. “Did v2 actually improve on v1, or just shift the failure modes?”
The typical workflow: instrument your application to send traces to Langfuse, attach automated scores, periodically run LLM-as-judge evaluations on a sample, and use the dashboards to spot trends.
LangSmith
LangSmith is LangChain’s commercial observability platform. Similar capabilities to Langfuse, with some differences:
- Cloud-hosted (no self-hosting option for the full platform)
- Tighter integration with LangChain/LangGraph
- Better UI for side-by-side run comparison
- Annotation queues for team-based human review workflows
- More mature dataset and experiment management
When to use which: LangSmith if you’re building with LangChain or need team-based annotation workflows. Langfuse if you prefer self-hosting and open source.
Other Tools
- Braintrust: Developer-focused eval platform with good dataset management and custom scorers.
- Arize Phoenix: Open-source, strong on embedding drift detection and retrieval analysis.
- Weights & Biases (W&B): If you use W&B for training metrics, W&B Prompts adds LLM tracing on top.
- OpenTelemetry: For infrastructure-level observability (latency, throughput, error rates). Complements LLM-specific tools for the “is the service healthy?” layer.
LLM-as-Judge
The most scalable evaluation technique: use a stronger model to evaluate a weaker model’s outputs.
How It Works
The idea is simple: take your fine-tuned model’s output, send it to a stronger model (like Claude or GPT-4) along with the original input and a rubric, and ask it to rate the quality. The judge model returns structured scores across your evaluation dimensions.
For QuizMe, a judge prompt might look like:
Given this question, correct answer, student answer, and the AI grader’s output, rate the grading quality on four dimensions (1-5 each): score accuracy, feedback quality, explanation accuracy, and internal consistency. Respond in JSON.
The judge sees the full context and rates whether the grader did a good job, not whether the student’s answer was correct (the grader already handled that).
Why This Works
- Scale: You can evaluate thousands of examples automatically, compared to dozens by human review.
- Consistency: Unlike human reviewers, the judge applies the same rubric every time (at temperature 0).
- Cost: Much cheaper than human evaluation, though not free (you’re paying for API calls to the judge model).
- Granularity: You get per-dimension scores, not just thumbs up/down.
Known Pitfalls
LLM-as-judge is powerful but has systematic biases you need to account for:
- Self-enhancement bias: Models rate their own family’s outputs higher. Use a different model family as judge (e.g., Claude judging LLaMA outputs).
- Verbosity bias: Judges tend to prefer longer answers, even when brevity is better. Explicitly instruct against this in your rubric.
- Position bias: In pairwise comparisons (“which response is better: A or B?”), judges prefer whichever is presented first. Randomize order and run both directions.
- Sycophancy: Judges may rate outputs as “good” even when they’re mediocre, especially with vague rubrics. Make your criteria specific and include examples of each score level.
Human Review Still Matters
LLM-as-judge is a filter, not a replacement for human judgment. A good cadence:
- Every inference: Automated format checks (JSON valid? all fields present?)
- Daily/weekly: LLM-as-judge on a random sample of production traces
- Monthly: Human review of 50-100 examples, focusing on cases where the judge gave borderline scores
Direct Preference Optimization (DPO)
DPO is a technique for the alignment stage (Stage 3 from Part 1). While supervised fine-tuning teaches the model what to output, DPO teaches it to prefer certain outputs over others.
When to Use DPO
DPO is useful when your SFT model works mostly but has specific failure modes you want to eliminate:
- The model defaults to giving “5/10” too often (safe, non-committal)
- Feedback is vague (“Good answer” vs. specific, actionable feedback)
- The model occasionally produces outputs in the wrong format
- Explanations are correct but not educational
How It Works
Instead of showing the model correct outputs (SFT), you show it pairs of outputs and tell it which one is better:
{
"prompt": "Grade this answer: ...",
"chosen": "Detailed, well-calibrated grading with specific feedback...",
"rejected": "Vague grading with generic feedback..."
}The DPO loss function directly adjusts the model’s weights to increase the probability of “chosen” responses and decrease the probability of “rejected” ones, relative to a reference model (usually your SFT checkpoint).
The math behind DPO (simplified): instead of training a separate reward model and using reinforcement learning (RLHF), DPO derives a closed-form solution that lets you optimize directly from preference pairs. The result is the same, but the training pipeline is much simpler.
Generating DPO Data
The most practical way to build a DPO dataset:
- Take your SFT model and generate multiple responses (5-10) for each input at moderate temperature
- Use LLM-as-judge to score each response
- Pair the best and worst responses as chosen/rejected
- Repeat across your eval dataset to build 200-500+ pairs
This creates a dataset of your model’s own outputs, ranked by quality. DPO then pushes the model toward its best behavior and away from its worst.
DPO vs. More SFT
A common question: why not just add more high-quality examples to your SFT dataset?
- SFT teaches “what good looks like” but doesn’t explicitly teach “what bad looks like.” The model can still produce bad outputs if they’re in a region of the distribution that SFT data didn’t cover.
- DPO explicitly contrasts good and bad, giving the model a signal about what to avoid. This is particularly effective for fixing specific failure modes.
In practice, do SFT first, evaluate, then use DPO to address remaining issues. DPO is a refinement tool, not a replacement for SFT.
Production Monitoring
Once your model is deployed, you need ongoing monitoring to catch degradation before users notice it.
Key Metrics
| Metric | What it catches | Alert threshold |
|---|---|---|
| Format compliance rate | Model producing malformed output | < 95% |
| Average judge score | Overall quality degradation | < 3.5/5.0 |
| Latency p95 | Performance regression | > 5s (API), > 15s (on-device) |
| Token usage per request | Model becoming verbose or looping | > 2x baseline |
| Score distribution entropy | Model collapsing to one score | Sudden narrowing |
| User-reported issues | Problems automated checks miss | Any spike |
The Continuous Improvement Loop
The goal isn’t “deploy and forget.” It’s a cycle:
Production traffic → Sample traces → Human review subset
↓ ↓
Monitor metrics Add to eval dataset
↓ ↓
Detect degradation Re-run evals on new model
↓ ↓
Alert / rollback ←──── Deploy if improved ←┘Collect
Sample 5-10% of production traces. Store inputs and outputs in your observability platform.
Evaluate
Weekly: run LLM-as-judge on the sample. Monthly: human review of 50-100 examples, prioritizing cases where the judge gave borderline or low scores.
Identify
Look for systematic errors, not one-off failures. Common patterns: specific subjects scored poorly, certain answer types mishandled, format regressions after a model update.
Improve
Add failing cases to the training set with corrected outputs. Re-fine-tune. Run the full eval suite to confirm improvement without regression on other dimensions.
Deploy
A/B test or canary deploy the new model version. Monitor metrics for 48 hours before full rollout.
Putting It All Together
The full development-to-production lifecycle:
┌──────────────────────────────────────────────────┐
│ Development │
│ MLX (training) → Eval datasets │
│ LLM-as-judge (quality) → W&B (training metrics) │
└──────────────────────┬───────────────────────────┘
│
┌──────────────────────▼───────────────────────────┐
│ Deployment │
│ CoreML (iOS) ←── merged model ──→ Ollama (API) │
│ Observability platform (traces + scores) │
└──────────────────────┬───────────────────────────┘
│
┌──────────────────────▼───────────────────────────┐
│ Continuous Improvement │
│ Sample traces → Judge → Identify gaps │
│ → Expand dataset → Re-train → Eval → Deploy │
└──────────────────────────────────────────────────┘The key insight: evaluation is not a one-time event. It’s an ongoing process that runs in parallel with production. The better your evaluation pipeline, the more confidently you can iterate on the model.
Further Reading
References
- Langfuse Documentation: Open-source LLM observability.
- LangSmith Documentation: LangChain’s observability platform.
- Judging LLM-as-a-Judge (Zheng et al., 2023): The paper that formalized LLM-as-judge evaluation.
- DPO: Direct Preference Optimization (Rafailov et al., 2023): The paper that simplified alignment training.
- LMSYS Chatbot Arena: Real-world LLM comparison via human voting. Good calibration reference.
- Anthropic’s Evals Guide: Practical evaluation methodology.