Skip to content

How Generation Works

The transformer processes your input and produces… what, exactly? Not text. It produces a vector of scores for every possible next token. This section covers how those scores become actual words, and the parameters you’ll use to control the output.

From Logits to Probabilities

After the input passes through all transformer layers, the final output is a vector of size vocab_size (typically 32K-128K) for each position in the sequence. These raw scores are called logits.

Each logit is an unnormalized score: a higher number means the model thinks that token is more likely to come next. But they’re not probabilities yet, they can be negative and don’t sum to 1.

To convert logits into a probability distribution, we apply softmax:

P(tokeni)=ezi∑jezjP(\text{token}_i) = \frac{e^{z_i}}{\sum_j e^{z_j}}

where ziz_i is the logit for token ii.

What softmax does:

  1. Raise ee (2.718…) to the power of each logit. This makes all values positive.
  2. Divide each by the sum of all of them. This makes them sum to 1.
  3. Larger logits become larger probabilities. The biggest logit dominates.

After softmax, you have a probability distribution over the entire vocabulary. For the sentence “The capital of France is,” the distribution might look like:

"Paris":  0.85
"Lyon":   0.03
"the":    0.02
"a":      0.01
...       (32K more tokens, all with tiny probabilities)

Sampling Strategies

The model doesn’t just pick the highest-probability token every time. That’s called greedy decoding, and it tends to produce repetitive, boring text (the model gets stuck in safe, common patterns).

Instead, the model samples from the probability distribution, with several knobs to control how “creative” or “focused” the output is.

Temperature

Temperature (TT) is the most important generation parameter. It reshapes the probability distribution before sampling:

P(tokeni)=ezi/T∑jezj/TP(\text{token}_i) = \frac{e^{z_i / T}}{\sum_j e^{z_j / T}}

The key insight: dividing logits by TT before softmax changes the “sharpness” of the distribution.

Temperature Effect Distribution shape
T=1.0T = 1.0 Default. No change. As the model learned it
T<1.0T < 1.0 (e.g., 0.3) Sharper. High-probability tokens get even higher. Low-probability tokens get crushed toward zero. More deterministic, “safer”
T>1.0T > 1.0 (e.g., 1.5) Flatter. Probabilities become more uniform. Low-probability tokens get a chance. More random, “creative”
T→0T \to 0 Equivalent to greedy decoding. Always picks the top token. Completely deterministic

For fine-tuning inference, you typically want low temperature (0.1-0.3) for consistent, predictable outputs (like JSON grading). Higher temperature (0.7-1.0) is better for creative tasks.

Top-k Sampling

Before sampling, zero out everything except the kk highest-probability tokens. This creates a hard cutoff: no matter how creative the temperature makes things, the model can never pick an extremely unlikely token.

k=50: Consider only the top 50 tokens, ignore the other 31,950+

Top-p (Nucleus Sampling)

Instead of a fixed count, keep the smallest set of tokens whose cumulative probability exceeds pp.

p=0.9: Keep tokens until their probabilities sum to 90%, discard the rest

This is adaptive: when the model is confident (one token has 95% probability), top-p considers very few alternatives. When it’s uncertain (probabilities spread across many tokens), it considers more. This generally produces better results than top-k.

Repetition Penalty

Reduce the logits of tokens that already appear in the generated text. This prevents the model from getting stuck in loops like “The cat sat on the mat. The cat sat on the mat. The cat…”

How These Interact

In practice, you combine these:

# Typical settings for a QuizMe grading model
temperature=0.3,    # Low: we want consistent grading
top_p=0.9,          # Moderate: allow some variety in wording
repetition_penalty=1.1  # Slight penalty to avoid repeated phrases
# Typical settings for creative writing
temperature=0.8,    # Higher: we want variety
top_p=0.95,         # Permissive
top_k=50,           # Safety net against wild tokens
Temperature annealing is a technique where you start generation with higher temperature (diverse, creative) and gradually reduce it (convergent, precise). Useful for generating varied training data that still stays on-topic.

Autoregressive Generation

LLMs generate text one token at a time:

Process the prompt

Feed the entire prompt through all transformer layers at once. This produces a probability distribution for the next token after the prompt.

Sample a token

Pick a token from the distribution (using temperature, top-k, top-p as described above).

Append and repeat

Add the sampled token to the input sequence. Run the model again to get the distribution for the next token. Repeat until a stop token is generated or a maximum length is reached.

Why Generation Is Slow

Prompt processing is fast because all tokens are processed in parallel (one matrix multiplication for the whole sequence). But generation is inherently sequential: you can’t predict token 5 until you know what token 4 is, because token 4 is part of the input for predicting token 5.

This means each new token requires a forward pass through the entire model. For a 3B model generating 200 tokens, that’s 200 sequential forward passes.

KV Caching

There’s an important optimization: KV caching. During generation, the Key and Value tensors for all previous tokens don’t change (because we’re only appending, never modifying). So the model caches them and only computes Q/K/V for the new token at each step.

Without KV cache: each generation step recomputes attention for all tokens (gets slower as the output grows). With KV cache: each generation step only computes attention for the new token against the cached K/V (constant time per step).

This is why you’ll see “KV cache” mentioned in model deployment configs, and why it matters for memory usage during inference.


Practical Numbers

To give you a sense of real-world generation speed on your MacBook Pro:

Model Size Tokens/sec Time for 200 tokens
LLaMA 3.2 1B (4-bit) ~0.5 GB ~40-60 t/s ~3-5 seconds
LLaMA 3.2 3B (4-bit) ~1.5 GB ~20-30 t/s ~7-10 seconds
LLaMA 3.1 8B (4-bit) ~4 GB ~10-15 t/s ~13-20 seconds

These are rough numbers for Apple Silicon via MLX. Actual speed depends on your specific chip, available memory, and sequence length.