How Generation Works
The transformer processes your input and produces… what, exactly? Not text. It produces a vector of scores for every possible next token. This section covers how those scores become actual words, and the parameters you’ll use to control the output.
From Logits to Probabilities
After the input passes through all transformer layers, the final output is a vector of size vocab_size (typically 32K-128K) for each position in the sequence. These raw scores are called logits.
Each logit is an unnormalized score: a higher number means the model thinks that token is more likely to come next. But they’re not probabilities yet, they can be negative and don’t sum to 1.
To convert logits into a probability distribution, we apply softmax:
where is the logit for token .
What softmax does:
- Raise (2.718…) to the power of each logit. This makes all values positive.
- Divide each by the sum of all of them. This makes them sum to 1.
- Larger logits become larger probabilities. The biggest logit dominates.
After softmax, you have a probability distribution over the entire vocabulary. For the sentence “The capital of France is,” the distribution might look like:
"Paris": 0.85
"Lyon": 0.03
"the": 0.02
"a": 0.01
... (32K more tokens, all with tiny probabilities)Sampling Strategies
The model doesn’t just pick the highest-probability token every time. That’s called greedy decoding, and it tends to produce repetitive, boring text (the model gets stuck in safe, common patterns).
Instead, the model samples from the probability distribution, with several knobs to control how “creative” or “focused” the output is.
Temperature
Temperature () is the most important generation parameter. It reshapes the probability distribution before sampling:
The key insight: dividing logits by before softmax changes the “sharpness” of the distribution.
| Temperature | Effect | Distribution shape |
|---|---|---|
| Default. No change. | As the model learned it | |
| (e.g., 0.3) | Sharper. High-probability tokens get even higher. Low-probability tokens get crushed toward zero. | More deterministic, “safer” |
| (e.g., 1.5) | Flatter. Probabilities become more uniform. Low-probability tokens get a chance. | More random, “creative” |
| Equivalent to greedy decoding. Always picks the top token. | Completely deterministic |
For fine-tuning inference, you typically want low temperature (0.1-0.3) for consistent, predictable outputs (like JSON grading). Higher temperature (0.7-1.0) is better for creative tasks.
Top-k Sampling
Before sampling, zero out everything except the highest-probability tokens. This creates a hard cutoff: no matter how creative the temperature makes things, the model can never pick an extremely unlikely token.
k=50: Consider only the top 50 tokens, ignore the other 31,950+Top-p (Nucleus Sampling)
Instead of a fixed count, keep the smallest set of tokens whose cumulative probability exceeds .
p=0.9: Keep tokens until their probabilities sum to 90%, discard the restThis is adaptive: when the model is confident (one token has 95% probability), top-p considers very few alternatives. When it’s uncertain (probabilities spread across many tokens), it considers more. This generally produces better results than top-k.
Repetition Penalty
Reduce the logits of tokens that already appear in the generated text. This prevents the model from getting stuck in loops like “The cat sat on the mat. The cat sat on the mat. The cat…”
How These Interact
In practice, you combine these:
# Typical settings for a QuizMe grading model
temperature=0.3, # Low: we want consistent grading
top_p=0.9, # Moderate: allow some variety in wording
repetition_penalty=1.1 # Slight penalty to avoid repeated phrases# Typical settings for creative writing
temperature=0.8, # Higher: we want variety
top_p=0.95, # Permissive
top_k=50, # Safety net against wild tokensAutoregressive Generation
LLMs generate text one token at a time:
Process the prompt
Feed the entire prompt through all transformer layers at once. This produces a probability distribution for the next token after the prompt.
Sample a token
Pick a token from the distribution (using temperature, top-k, top-p as described above).
Append and repeat
Add the sampled token to the input sequence. Run the model again to get the distribution for the next token. Repeat until a stop token is generated or a maximum length is reached.
Why Generation Is Slow
Prompt processing is fast because all tokens are processed in parallel (one matrix multiplication for the whole sequence). But generation is inherently sequential: you can’t predict token 5 until you know what token 4 is, because token 4 is part of the input for predicting token 5.
This means each new token requires a forward pass through the entire model. For a 3B model generating 200 tokens, that’s 200 sequential forward passes.
KV Caching
There’s an important optimization: KV caching. During generation, the Key and Value tensors for all previous tokens don’t change (because we’re only appending, never modifying). So the model caches them and only computes Q/K/V for the new token at each step.
Without KV cache: each generation step recomputes attention for all tokens (gets slower as the output grows). With KV cache: each generation step only computes attention for the new token against the cached K/V (constant time per step).
This is why you’ll see “KV cache” mentioned in model deployment configs, and why it matters for memory usage during inference.
Practical Numbers
To give you a sense of real-world generation speed on your MacBook Pro:
| Model | Size | Tokens/sec | Time for 200 tokens |
|---|---|---|---|
| LLaMA 3.2 1B (4-bit) | ~0.5 GB | ~40-60 t/s | ~3-5 seconds |
| LLaMA 3.2 3B (4-bit) | ~1.5 GB | ~20-30 t/s | ~7-10 seconds |
| LLaMA 3.1 8B (4-bit) | ~4 GB | ~10-15 t/s | ~13-20 seconds |
These are rough numbers for Apple Silicon via MLX. Actual speed depends on your specific chip, available memory, and sequence length.