Skip to content

Tokens & Embeddings

Before an LLM can do anything with text, it needs to convert that text into numbers. This happens in two steps: tokenization (text to integer IDs) and embedding (integer IDs to vectors). These steps are simple but they have consequences that matter when you fine-tune.

Tokenization

LLMs don’t see characters or words. They see tokens: integer IDs from a fixed vocabulary. A tokenizer converts text into a sequence of these IDs and back.

"Hello world" → [9906, 1917]

How Vocabularies Are Built

Most modern models use Byte-Pair Encoding (BPE) or a variant called SentencePiece. The algorithm works like this:

  1. Start with every individual byte (or character) as its own token
  2. Scan the training corpus and find the most frequent pair of adjacent tokens
  3. Merge that pair into a new, single token
  4. Repeat until you hit a target vocabulary size (typically 32K-128K tokens)

The result is that common words and phrases become single tokens, while rare words get split into pieces (example with the tokenizer for LLaMA 3 8B):

"the"          → [1820]              # Common word = 1 token
"Unsupervised" → ["Un", "super", "vised"] → [1844, 13066, 79090]  # 3 tokens
"PyTorch"      → ["Py", "T", "orch"]      → [14149, 51, 22312]   # 3 tokens

This is why token count is not the same as word count. A short technical term might be 3 tokens while “the” is always 1.

Why Tokenization Matters for Fine-Tuning

Tokenization is fixed for a given model. When you fine-tune, you’re adjusting the model’s weights, not its vocabulary. You cannot add new tokens or change how existing text gets split.

This has practical consequences:

  • Domain-specific terminology (medical codes, API names, non-English words) may get split into many small tokens. The model sees fragments, not whole concepts. If your fine-tuning data is full of these, the model has to work harder to learn them.
  • Cost scales with tokens, not words. If your domain text tokenizes inefficiently (3-4 tokens per word instead of 1-2), you’ll pay more for API inference and need longer sequence lengths during training.
  • All models in the same family share a tokenizer. LLaMA 3.2 1B and LLaMA 3.1 70B use the same vocabulary, so tokenization behavior is identical across sizes.

You can inspect any model’s tokenizer interactively at Tiktokenizer to see how your data gets split.


Embeddings

Once text is tokenized into integer IDs, the next step is converting those IDs into something a neural network can work with: vectors.

What Is an Embedding?

Each token ID maps to a dense vector of floating-point numbers. Think of it as a list of coordinates that describe where that token sits in a high-dimensional space.

For a model with embedding dimension d_model = 4096, each token becomes a 4096-dimensional vector:

token 9906 ("Hello") → [0.023, -0.841, 0.117, ..., 0.502]  (4096 floats)
token 1917 ("world") → [0.195, -0.230, 0.884, ..., -0.112]  (4096 floats)

These 4096 numbers aren’t hand-crafted. They’re learned during training. As the model sees billions of sentences, it adjusts these vectors so that tokens appearing in similar contexts end up pointing in similar directions.

The Geometry of Meaning

This learned structure has a useful property: semantic relationships become geometric relationships.

  • “King” and “queen” end up close together (both are royalty)
  • “King” and “toaster” end up far apart (unrelated concepts)
  • The classic example: vector("king") - vector("man") + vector("woman") ≈ vector("queen")

This isn’t a parlor trick. It means the model has an internal representation of meaning, encoded as directions and distances in this high-dimensional space. When you fine-tune, you’re adjusting these representations (indirectly, through the model’s weights) to better capture the patterns in your training data.

Positional Encoding

One problem: the embedding for “Hello” is the same regardless of whether it appears at position 1 or position 500 in the input. But word order obviously matters (“dog bites man” vs. “man bites dog”).

Positional encoding solves this by adding position information to each embedding. Modern models (LLaMA, Mistral) use Rotary Position Embeddings (RoPE), which encode relative position through rotation in the embedding space. The details aren’t critical for fine-tuning, but the key consequence is:

  • Models can generalize to sequence lengths longer than they were trained on (to a degree)
  • The model always knows the relative ordering of tokens

The Full Picture So Far

At this point, we’ve gone from raw text to a matrix of vectors:

"Hello world" → tokenize → [9906, 1917] → embed → [[0.023, -0.841, ...],
                                                  [0.195, -0.230, ...]]
                                                  (2 × 4096 matrix)

Each row is a token, and each column is one dimension of the embedding space. This matrix is the input to the transformer, which we’ll cover next.


Further Reading

References
  • Tiktokenizer: Interactive tokenizer visualization. Paste your text and see how different models split it.
  • The Illustrated Word2Vec: Jay Alammar’s visual guide to word embeddings. The concepts carry over to LLM token embeddings.
  • Hugging Face Tokenizer Docs: Deep dive into BPE, WordPiece, and SentencePiece implementations.