Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Tokenization Deep Dive: How Text Becomes Model Input

By Anup Rai31 min readReviewed September 2026

Text tokenization segments text into discrete units called tokens. For a language model, the tokenizer also maps those units to integer token IDs in its vocabulary. A decoder reconstructs text from IDs according to the same tokenizer configuration.

Tokenization affects:

  • how much text fits in a context window,
  • how much an API request costs,
  • which spelling, code, and multilingual patterns are easy or awkward for the model,
  • how text chunks line up with user-visible characters,
  • and whether a prompt is serialized in the format the model saw during training.

This chapter zooms into the token step introduced in LLM Fundamentals. It does not repeat embeddings or Transformer layers. After text becomes token IDs here, chapter 01 explains how IDs select embeddings, Attention Mechanisms explains how positions exchange information, and Transformer Architecture assembles the full model.


Table of Contents

  1. The whole job in one picture
  2. Tokens, IDs, and text spans
  3. Why subwords occupy the useful middle
  4. The complete tokenizer pipeline
  5. Byte Pair Encoding
  6. WordPiece
  7. Unigram and SentencePiece
  8. Choosing among subword algorithms
  9. Vocabulary-size tradeoffs
  10. Unicode, normalization, and byte fallback
  11. Whitespace and word-boundary markers
  12. Special tokens and chat templates
  13. Multilingual and domain tokenization
  14. Multimodal tokenization
  15. Counting and budgeting tokens
  16. Chunking without corrupting text
  17. Common failure modes
  18. How to evaluate a tokenizer
  19. Interview questions
  20. Compact reference
  21. Engineering references

1. The Whole Job in One Picture

Tokenizer pipeline: raw text moves through normalization, pre-tokenization, a learned subword model, and post-processing before becoming token IDs; decoding reverses IDs into text pieces.

For the text:

The robot waved.

one tokenizer might produce pieces like:

["The", " robot", " waved", "."]

and map them to IDs such as:

[791, 12585, 23405, 13]

Those numbers are illustrative. Another tokenizer can split the same text differently and assign entirely different IDs.

The model does not attach meaning to the size of an ID. Token ID 23405 is not “more meaningful” than ID 13. Each ID is an address into a vocabulary and, from there, an embedding row.

Encoding and decoding

Encoding usually returns more than IDs:

text
  → normalized text
  → token pieces
  → token IDs
  → offsets / attention mask / special-token metadata

Decoding maps IDs back to pieces and then joins those pieces according to the tokenizer's rules:

token IDs
  → token pieces
  → reconstructed text

Decoding is not always a perfect inverse of the original string. A normalizer might lowercase text, replace Unicode forms, or remove information before tokenization. The pipeline can be reversible only with respect to what it preserved.


2. Tokens, IDs, and Text Spans

Three objects are easy to mix up:

Object Example What it is
Text span characters 4 through 9 A region in the original or normalized string
Token piece " robot" The visible vocabulary piece produced for that span
Token ID 12585 The integer vocabulary index consumed by the model

A token is not necessarily a word

A token can represent:

  • a complete common word,
  • part of a word,
  • whitespace plus a word fragment,
  • punctuation,
  • one byte or several bytes,
  • a control marker that never appeared in the user's visible text,
  • or a special marker associated with a modality boundary.

For a discrete text tokenizer, one token ID selects one vocabulary entry. Multimodal systems also use “token” for continuous image or audio feature vectors at sequence positions; those need not have discrete vocabulary IDs. Distinguish these meanings when discussing architecture and billing.

Offsets connect tokens back to text

An encoding can preserve an offset for every token:

text:    "red fox"
token:   "red"      " fox"
offset:  [0, 3)     [3, 7)

Offsets matter for:

  • highlighting search results,
  • mapping named entities to source text,
  • redacting sensitive spans,
  • attaching citations,
  • and splitting documents without cutting a token in the middle.

Check both the coordinate system and the unit: original versus normalized string, UTF-8 bytes versus Unicode code points, and browser UTF-16 code units. Normalization can change length, and some tokens share source spans or have special-token sentinel offsets. For example, Python counts 🙂 as one code point; JavaScript string length counts two UTF-16 code units. Convert offsets explicitly before highlighting or redacting text.


3. Why Subwords Occupy the Useful Middle

Suppose the model must represent unbelievable.

Word, subword, character, and byte tokenization comparison showing the tradeoff between vocabulary size and sequence length.

One token per word

A word vocabulary gives short sequences for known words, but it grows without bound:

believe
believes
believed
believing
unbelievable

Names, typos, product codes, and newly coined words require more entries. Anything absent from the vocabulary may collapse into an unknown token, losing its internal spelling.

One token per character

A character vocabulary can spell unseen words whose characters are covered by its alphabet. Full Unicode coverage is much larger than an English alphabet; fallback still matters. The price is long sequences. The model must spend several attention positions rebuilding every common word from tiny pieces.

One token per byte

A byte vocabulary can represent any byte sequence with a fixed base set of 256 values. It avoids an unknown-token dead end, but familiar text may require many positions unless the tokenizer also learns larger byte sequences.

Subword tokens

Subword methods learn that common sequences deserve shortcuts while rare sequences can be spelled from smaller pieces:

un + believ + able

This gives a useful compromise:

  • frequent text uses fewer positions,
  • uncommon text remains representable when its characters or bytes are covered,
  • and the vocabulary stays much smaller than a pure word vocabulary.

This is why BPE, WordPiece, and Unigram appear throughout modern language-model tooling.


4. The Complete Tokenizer Pipeline

Calling a tokenizer can look like one function call, but several policies run in sequence.

4.1 Normalization

Normalization transforms the input before pieces are selected. Possible operations include:

  • Unicode normalization such as NFC or NFKC,
  • lowercasing,
  • accent handling,
  • whitespace cleanup,
  • or application-specific character replacement.

There is no universal “clean text” rule. Lowercasing may be useful for an uncased retrieval model and destructive for code, names, chemical notation, or case-sensitive identifiers.

4.2 Pre-tokenization

A pre-tokenizer establishes candidate boundaries before the learned subword algorithm runs. It might split around:

  • spaces,
  • punctuation,
  • digits,
  • script changes,
  • or byte-level boundary rules.

Think of pre-tokenization as fencing the search area. If a boundary forbids merging across whitespace, the learned model cannot later create a token that crosses that fence.

4.3 The learned model

This is where BPE, WordPiece, Unigram, or a word-level model selects pieces and maps them to vocabulary IDs.

The tokenizer model here is not the neural language model. It is a smaller learned or derived segmentation model stored with the vocabulary and rules.

4.4 Post-processing

Tokenization post-processing can add control IDs and segment metadata to an encoding:

[CLS] sentence A [SEP] sentence B [SEP]

Chat templating is a related but separate step: it commonly formats structured messages into text before tokenization. A low-level tokenizer post-processor and a chat template are not interchangeable. Apply the model's documented pipeline exactly once.

4.5 Truncation and padding

Batching may require:

  • truncation to a maximum token length,
  • padding shorter examples to a shared tensor length,
  • and an attention mask so the model ignores padding positions.

Truncation is an information policy, not a harmless tensor operation. “Keep the first 8,000 tokens” may remove the conclusion, the current question, or the last tool result.


5. Byte Pair Encoding

Byte Pair Encoding (BPE) learns larger pieces by repeatedly merging selected adjacent pairs.

BPE merge ladder for the word lower: characters combine into lo, low, and larger reusable pieces as adjacent pairs are learned.

Training intuition

Start from small units across a training corpus. Depending on the implementation, the base units may be characters, encoded symbols, or bytes.

Then repeat:

  1. count adjacent pairs,
  2. select a frequent pair according to the training rule,
  3. replace that pair with one new vocabulary piece,
  4. record the merge,
  5. stop when the vocabulary budget is reached.

Toy corpus:

low lower lowest

Early states might look like:

l o w
l o w e r
l o w e s t

lo w
lo w e r
lo w e s t

low
low e r
low e s t

Because l + o and then lo + w occur repeatedly, low becomes a useful piece.

Encoding with trained merges

Training creates the vocabulary and merge priority. Encoding new text does not retrain BPE. Ordinary inference applies the fixed merge ranks deterministically. Optional BPE dropout or sampled segmentation deliberately introduces variation; disable that mode for reproducible counting.

An unseen word can still be represented from smaller pieces:

lowest-ish → low + est + - + ish

The exact result depends on the learned corpus, normalization, pre-tokenization, and base alphabet.

BPE does not always mean byte-level BPE

The name describes the merge strategy. A BPE tokenizer can start from characters or from byte-derived symbols. Byte-level BPE specifically begins with byte coverage, which provides a route for arbitrary input bytes without requiring a normal unknown token.


6. WordPiece

WordPiece is another learned subword method, strongly associated with BERT-style tokenizers.

The useful contrast with BPE

At a high level:

  • BPE training chooses frequent merges.
  • WordPiece-style training scores how useful candidate pieces are to the corpus likelihood rather than using raw pair frequency alone.
  • Common WordPiece encoders then choose the longest vocabulary match at each point in a word.

Example vocabulary:

play
##ing
##ful

Possible encoding:

playful → play + ##ful
playing → play + ##ing

The ## convention says that the piece continues a word. It is a display convention in many WordPiece vocabularies, not a universal symbol used by every tokenizer.

Greedy longest-match encoding

For one pre-tokenized word:

  1. begin at the first character,
  2. find the longest vocabulary piece that matches,
  3. emit its ID,
  4. continue from the unmatched suffix,
  5. use an unknown-token policy if no valid decomposition exists.

This encoder is deterministic once the vocabulary and rules are fixed.

Distinguish the training description from the deployed encoder

The original Google WordPiece trainer was not released. The frequently taught score freq(a,b) / (freq(a) × freq(b)) is a reconstruction used to explain likelihood-oriented selection, not a universal training contract. Different trainers can build a WordPiece vocabulary differently. The deployed BERT-style encoder uses longest-match segmentation, commonly returning [UNK] for an entire pre-tokenized word when it cannot complete the decomposition. See the Hugging Face WordPiece explanation and caveat.


7. Unigram and SentencePiece

Unigram approaches segmentation from the opposite direction.

Unigram intuition

Start with a large set of candidate pieces. Assign probabilities to them. Then repeatedly remove pieces whose loss has the smallest harmful effect, retraining the remaining probabilities as the vocabulary shrinks.

For one string, several segmentations may be possible:

un + believable
un + believ + able
u + n + believ + able

The model scores whole segmentations and chooses a likely path. A Viterbi-style dynamic program can find the best path efficiently.

Because multiple paths have probabilities, Unigram tokenizers can also sample alternative segmentations during training. That technique can act as data augmentation, often called subword regularization.

SentencePiece is a toolkit, not one algorithm

This distinction is routinely missed:

SentencePiece can train BPE or Unigram tokenizers. “SentencePiece” and “Unigram” are not synonyms.

SentencePiece treats input as a raw character stream rather than requiring a language-specific word splitter first. A visible marker such as ▁ can represent a preceding space:

"hello world" → ["▁hello", "▁world"]

The marker preserves spaces in the normalized representation. It does not guarantee byte-for-byte recovery of the original string. Default SentencePiece normalization can fold Unicode forms and collapse or strip whitespace; configure and test the required preservation behavior. See SentencePiece normalization. No language-specific word splitter is required before its raw-string model.


8. Choosing Among Subword Algorithms

Comparison of BPE building upward through frequent merges, WordPiece selecting useful pieces and using longest match, and Unigram pruning a large candidate vocabulary.

Method Training picture Common encoding picture Useful property Watch for
BPE Begin small; merge selected adjacent pairs Replay merge priorities Simple, deterministic, widely implemented “BPE” does not imply byte coverage
WordPiece Select pieces with a likelihood-oriented score Greedy longest match Strong established ecosystem Unknown-token behavior and implementation details
Unigram Begin large; prune weak pieces Choose highest-probability segmentation Multiple candidate segmentations More probabilistic machinery
SentencePiece Toolkit that can train BPE or Unigram Encodes raw strings with its chosen model Raw-string processing and explicit normalized-space markers Do not call it a fourth segmentation algorithm

No method is automatically best for every model. Corpus design, normalization, byte fallback, vocabulary size, and downstream data can matter as much as the algorithm name.

What must travel with model weights?

A model and tokenizer form a contract. See embedding compatibility for the corresponding vector-space boundary. A deployable tokenizer package commonly needs:

  • the vocabulary-to-ID mapping,
  • merge rules or model probabilities,
  • normalizer configuration,
  • pre-tokenizer rules,
  • special-token IDs,
  • post-processing or chat-template rules,
  • decoder rules,
  • and a version identifier.

Using “roughly the same” tokenizer with a checkpoint is not enough. If token IDs change, embedding rows point to the wrong learned vectors.


9. Vocabulary-Size Tradeoffs

Graph showing that larger vocabularies tend to reduce sequence length while increasing embedding and output-table memory, with a workload-dependent useful middle.

Suppose the model hidden width is d_model and the vocabulary contains V entries. The input embedding table has roughly:

V × d_model parameters

The language-model output projection also maps to V logits. Some architectures tie its weights to the input embedding table; others store a separate table.

Larger vocabulary

Possible benefits:

  • frequent phrases need fewer tokens,
  • sequence length decreases on the target distribution,
  • more domain terms can remain intact.

Costs:

  • larger embedding and output layers,
  • a more expensive final vocabulary projection,
  • more rare entries that receive limited training,
  • and potentially less sharing among related word forms.

Smaller vocabulary

Possible benefits:

  • smaller model tables,
  • broad sharing of common fragments,
  • fewer extremely rare entries.

Costs:

  • longer sequences,
  • more attention positions,
  • higher prompt cost for the same text,
  • and more steps to generate the same visible answer.

Calculate the table-memory tradeoff

With hidden width 4,096 and BF16 parameters:

Vocabulary entries One input embedding table Untied input plus output tables
32,000 262,144,000 bytes = 250 MiB 500 MiB
128,000 1,048,576,000 bytes = 1,000 MiB 2,000 MiB

The larger vocabulary adds 750 MiB to one table, or 1,500 MiB if both are separate. Weight tying avoids a duplicate table; it does not eliminate the vocabulary projection. These calculations exclude biases, optimizer state and quantization metadata. The chart shows qualitative trends, not a measured universal optimum or guaranteed compression improvement.

The right measurement is workload-specific

Do not compare tokenizers only on English prose. Measure the traffic you expect:

  • supported languages,
  • source code,
  • JSON and SQL,
  • product identifiers,
  • mathematical notation,
  • URLs,
  • logs,
  • and user-generated spelling variation.

10. Unicode, Normalization, and Byte Fallback

Unicode assigns abstract code points to text characters. UTF-8 stores those code points as one or more bytes. A tokenizer can operate on characters, byte-derived symbols, or a mixture of learned larger pieces and byte fallback.

Unicode and UTF-8 relationship: examples map to one or more hexadecimal bytes, while byte fallback guarantees a representation for unseen input.

Visually identical text may have different code points

An accented character can be represented as:

  • one precomposed code point, or
  • a base character followed by a combining accent.

Unicode normalization can make these forms consistent. But a normalizer is part of the tokenizer contract. Changing it after training changes the ID sequence the model sees.

Byte fallback

Suppose a rare character has no learned piece. A tokenizer with byte fallback can encode its UTF-8 bytes instead of returning one undifferentiated unknown token.

With a complete byte fallback, the preserved valid UTF-8 input has a representation. This does not reverse earlier normalization or guarantee support for malformed raw bytes accepted by some other interface. It also does not make all input equally efficient. An unfamiliar script or noisy text may expand into many byte tokens.

“Character count” has several meanings

For user-visible text, distinguish:

  • bytes,
  • Unicode code points,
  • grapheme clusters perceived as one character,
  • and tokenizer pieces.

An emoji with a skin-tone modifier or a family joined by zero-width joiners can contain several code points while appearing as one grapheme. Tokenization may split it further. For user-perceived character boundaries, use a library implementing Unicode text segmentation, UAX #29, with its Unicode version recorded.

That is one reason an LLM can be unreliable at exact letter or character counting. Its main internal units are learned token positions, not a guaranteed array of user-visible graphemes.


11. Whitespace and Word-Boundary Markers

Whitespace may be:

  • attached to the following token,
  • attached to the previous token,
  • represented by an explicit marker,
  • normalized,
  • or encoded as bytes.

For example, a vocabulary might contain both:

"hello"
" hello"

Those are different pieces with different IDs.

Why leading spaces matter

Compare:

"hello"
" hello"

One appears at the start of a string; the other follows preceding content. If a tokenizer builds word-boundary information into tokens, the two contexts can produce different encodings.

Repeated whitespace

Indentation, tabs, and runs of spaces are common in code and tables. A tokenizer trained mostly on prose may encode them inefficiently or differently from a code-focused tokenizer.

Never trim or reformat prompt text merely to “clean it up” unless the application accepts the semantic change. Whitespace can carry structure in Python, Markdown, YAML, diffs, and fixed-width data.


12. Special Tokens and Chat Templates

Special tokens communicate structure rather than ordinary visible text.

Common roles include:

Role Example convention Purpose
Begin/end BOS, EOS Mark sequence boundaries or stopping
Padding PAD Fill a batch tensor to equal length
Unknown UNK Represent an unsupported span when fallback is unavailable
Classification CLS Provide a pooled position for some encoder models
Separation SEP Divide sentence segments
Masking MASK Hide a token during masked-language training
Chat roles system/user/assistant markers Serialize turns and instructions
Tool protocol tool call/result markers Delimit structured tool interactions

The strings and IDs are model-specific. There is no universal <eos> ID shared by all models.

Chat text is serialized text

Chat template serialization showing role markers, visible message text, assistant boundary, and end marker in the actual token sequence.

A chat UI may display:

System: Be concise.
User: Why is the sky blue?

The model might actually receive a serialized form conceptually like:

<system>
Be concise.
<user>
Why is the sky blue?
<assistant>

The exact template must match the model's training convention. The last assistant marker can tell the model whose turn comes next.

Why template mistakes hurt

Common errors include:

  • omitting role boundaries,
  • applying a template twice,
  • using one model family's markers with another checkpoint,
  • forgetting tool-result delimiters,
  • or counting visible message text but not hidden wrappers.

The output can still look superficially plausible while instruction following and tool behavior quietly degrade.

For a local Hugging Face chat checkpoint, apply_chat_template(..., tokenize=True) handles template serialization and tokenization together. If you format with tokenize=False and tokenize the resulting string later, use add_special_tokens=False to avoid duplicate BOS/EOS markers. add_generation_prompt=True requests a new assistant turn where the template supports it; continuing a partly written assistant message is a different operation. See the current chat-template contract.

User text is not an authorized control channel

A user can paste a string that resembles a special marker. Do not enable every special token or concatenate unescaped user content into a home-made control format. In tiktoken, ordinary text can be encoded with encode_ordinary; encode has explicit allowed/disallowed-special policies and may reject marker-shaped strings. This controls token recognition, not prompt-injection safety. Model roles and application authorization still come from trusted structured requests. See tiktoken's encoding contract.


13. Multilingual and Domain Tokenization

A shared vocabulary is learned from a finite corpus. Languages and domains that occupy more of that corpus often receive more efficient pieces.

Evaluation workflow for English, Hindi, Japanese and code samples, measuring token counts, coverage and task quality without inventing language-specific ratios.

Fertility

Tokenizer fertility commonly means the average number of subword tokens per word. State the word-segmentation convention. Tokens per byte, code point or grapheme are other useful efficiency ratios; label their denominators rather than calling them all the same measure. For an example of the standard per-word usage, see TokLens (ACL 2026).

If the same meaning takes 20 tokens in one language and 35 in another, then under the same token limit the second representation gets:

  • less visible text in context,
  • more attention positions,
  • potentially higher cost,
  • and more decode steps for an answer of similar visible length.

Fertility alone does not measure model quality. It is an efficiency and representation diagnostic.

Languages without spaces

Splitting on whitespace is not a language-independent definition of a word. Chinese, Japanese, Thai, and many mixed-script inputs require other boundary choices. SentencePiece-style raw-string modeling is useful partly because it does not demand a language-specific word segmenter first.

Code and structured data

Code-focused tokenizers may learn useful pieces for:

def
async
_id

": "
</

But token efficiency is not the only goal. Exact whitespace, quoting, and delimiter behavior must survive round trips.

Numbers and identifiers

A number may be one token, several digit groups, or individual digits. The same is true for UUIDs, hashes, and product IDs. Do not assume numerical magnitude is reflected in a token ID or one embedding.

For exact arithmetic, validation, and identifiers, use tools and parsers rather than relying on tokenization to preserve an ideal mathematical representation.


14. Multimodal Tokenization

“Token” generalizes beyond text, but the front end changes by modality.

Images

A Vision Transformer can divide an image into fixed-size patches. Each patch becomes a vector and one sequence position after projection. Other systems use learned visual encoders or discrete visual codebooks.

If a 224 × 224 image uses 16 × 16 non-overlapping patches:

14 patches per side × 14 patches per side = 196 patch positions

Extra class, separator, or image-boundary tokens may be added.

Audio

Audio front ends can create positions from:

  • spectrogram frames,
  • learned encoder frames,
  • or discrete codec codes.

Video

Video tokenization must handle space and time. Positions can come from frame patches, tubelets spanning multiple frames, or compressed learned codes.

The shared idea

Across modalities:

raw signal → manageable learned units → vectors → sequence model

Do not assume a visual or audio token corresponds to one human word or one discrete vocabulary entry. Architecture positions and billable tokens can differ: image resizing, crops, patch merging, audio compression and provider pricing rules affect the mapping. Use the chosen model's processor and counting endpoint rather than deriving an invoice from patch count alone.


15. Counting and Budgeting Tokens

Words are acceptable for a napkin estimate and unsafe for enforcement.

Count with the exact target tokenizer

The reliable procedure is:

  1. choose the exact model/tokenizer version,
  2. serialize the complete request, including role and tool wrappers,
  3. encode it,
  4. count the returned IDs,
  5. reserve room for the desired output and provider-specific limits.

For inspecting a particular text encoding locally:

import tiktoken

encoding = tiktoken.get_encoding("o200k_base")
token_ids = encoding.encode_ordinary("Token count this exact string.")
print(len(token_ids))

o200k_base names one encoding; it is not a claim that every current GPT model uses it. The snippet counts ordinary text only, not an entire hosted request.

Serving path, checked September 24, 2026 Counting method Boundary to preserve
OpenAI Responses API POST /v1/responses/input_tokens; Python client.responses.input_tokens.count(...) Include instructions, tools and conversation/input state accepted by the endpoint
Claude Messages API POST /v1/messages/count_tokens; Python client.messages.count_tokens(...) Include supported system, message, tool and media fields; count is an estimate
Gemini API models.countTokens, including its supported generateContentRequest representation Use the selected model and full supported request; bare text omits wrappers/configuration
Locally served checkpoint Its versioned tokenizer/processor and actual chat template Pin model, tokenizer, template and processor together

Use the provider's current schema, then reconcile estimates with reported usage after execution. A count call does not predict output tokens, prove a cache hit or reserve provider capacity. Sources: OpenAI counting, Claude counting, Gemini counting.

For an already downloaded, reviewed chat checkpoint, count a complete conversation locally:

from transformers import AutoTokenizer

# Replace with the reviewed local snapshot directory used by your model server.
tokenizer = AutoTokenizer.from_pretrained(
    "/path/to/pinned-chat-checkpoint", local_files_only=True,
    trust_remote_code=False,
)
messages = [{"role": "user", "content": "Explain tokenization in one sentence."}]
ids = tokenizer.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
)
print(len(ids))

Cost calculation

If a provider bills input and output separately:

request cost
  = input_tokens  × input_price_per_token
  + output_tokens × output_price_per_token

Use the provider's current prices and billing units. Cached input, reasoning tokens, tool calls, images, and batch processing may have separate rules.

Budget the whole context

Context-window budget divided among system and tool instructions, conversation history, retrieved evidence, user input, and reserved output.

For a context limit C:

available_for_dynamic_input
  = C
  - system_and_tool_tokens
  - reserved_output_tokens
  - safety_margin

Then allocate the remainder among history, retrieval, and the current user input based on information value.

Reserve output before filling input

If the context limit includes input plus output, filling every position with prompt tokens leaves no room for a useful answer. Reserve the answer budget first.

Keep a safety margin

Margins absorb:

  • hidden or versioned wrappers,
  • small counting differences,
  • tool schemas,
  • and application metadata added after initial planning.

The margin should be measured from real requests, not copied as a universal percentage.


16. Chunking Without Corrupting Text

Chunking for RAG or summarization has two goals:

  1. stay within a token budget,
  2. preserve meaningful boundaries.

Those goals can conflict.

Bad approach: slice characters by a token estimate

chunk = text[:4000]  # 4000 characters is not 4000 tokens

This can split:

  • a grapheme cluster,
  • a word,
  • a Markdown code fence,
  • a JSON string,
  • or the middle of the most important paragraph.

Better approach: structure first, tokens second

  1. parse the document into headings, paragraphs, list items, code blocks, or sentences,
  2. count each unit with the target tokenizer,
  3. combine adjacent units until reaching the budget,
  4. split an oversized unit with a smaller boundary rule,
  5. preserve source offsets and metadata,
  6. optionally add measured overlap.

Reference implementation for pre-split text units:

def pack_units(units, count_text, budget, separator="\n\n"):
    """Preserve complete text units; reject oversized ones for explicit splitting."""
    if type(budget) is not int or budget <= 0:
        raise ValueError("budget must be a positive integer")
    if not isinstance(separator, str):
        raise ValueError("separator must be text")

    def fits(text):
        size = count_text(text)
        if type(size) is not int or size < 0:
            raise ValueError("count_text must return a non-negative integer")
        return size <= budget

    chunks, current = [], []
    for unit in units:
        if not isinstance(unit, str) or not unit:
            raise ValueError("each unit must be non-empty text")
        if not fits(unit):
            raise ValueError("split an oversized unit before packing")
        candidate = separator.join(current + [unit])
        if current and not fits(candidate):
            chunks.append(separator.join(current))
            current = [unit]
        else:
            current.append(unit)
    if current:
        chunks.append(separator.join(current))
    return chunks

count_text must use the exact destination tokenizer with no implicit truncation. Every joined candidate is counted again, including separators: token counts need not be additive across text boundaries. This example rejects oversized units instead of silently returning an oversized chunk. Split them at paragraphs, sentences or grapheme boundaries, recount, and preserve source offsets in the enclosing application. The simple implementation can retokenize repeatedly; bound document size and benchmark it before optimizing.

A token boundary is not always a Unicode boundary

Byte-based token pieces can end in the middle of a UTF-8 character. Decoding arbitrary token slices independently may insert replacement characters. Preserve original source text and offsets, or join token bytes before a strict/incremental UTF-8 decoder. When streaming, keep decoder state across pieces and finalize it at the end; do not drop an incomplete final sequence silently. Valid UTF-8 boundaries still do not guarantee intact grapheme clusters or meaningful document structure.

Overlap is not free

Overlap can preserve context across boundaries, but it:

  • increases indexing and prompt tokens,
  • creates near-duplicate retrieval results,
  • and can crowd out more diverse evidence.

Choose overlap from retrieval evaluation rather than habit.


17. Common Failure Modes

Mistake 1: assuming one token equals one word

This breaks cost estimates, context checks, and explanations of spelling tasks.

Mistake 2: counting with the wrong tokenizer

The same text can have different boundaries and counts across model families or tokenizer revisions. Count with the exact deployed configuration.

Mistake 3: changing normalization without retraining

If the checkpoint learned from one ID distribution and production emits another, the model receives unfamiliar sequences.

Mistake 4: changing the vocabulary but keeping model weights

Vocabulary ID 500 points to a specific learned embedding row. Reordering IDs requires the same permutation of input embeddings, output rows/biases and all ID-dependent metadata. A consistent permutation alone need not require retraining; newly added pieces need appropriately trained weights. Changing segmentation itself is a broader model adaptation.

Mistake 5: double-applying a chat template

If a client library already serializes roles and the application adds its own markers, the model sees duplicated control text.

Mistake 6: decoding one token at a time for streaming

Some byte or Unicode sequences become valid text only after multiple token pieces are combined. A streaming decoder should preserve incremental decoder state rather than assuming every token ID is an independent printable string.

Mistake 7: treating byte fallback as equal language efficiency

Byte fallback guarantees coverage. It does not guarantee short sequences, well-trained representations, or equal quality across scripts.

Mistake 8: truncating from one end blindly

Keeping only the beginning may discard the current question. Keeping only the end may discard instructions. Use a component-aware policy.

Mistake 9: believing tokens explain all model errors

Tokenization can contribute to spelling and boundary problems, but reasoning errors also come from training data, architecture, decoding, context, and task difficulty. Inspect the actual encoding before blaming it.


18. How to Evaluate a Tokenizer

Use held-out, representative samples. When training a tokenizer, keep evaluation text separate from the training corpus; when evaluating a deployed tokenizer, use its fixed production configuration.

Coverage

  • What fraction requires an unknown token?
  • Does byte fallback cover noisy or unseen scripts?
  • Do encode/decode round trips preserve required text?

Efficiency

  • tokens per byte,
  • tokens per Unicode character or grapheme,
  • tokens per word where “word” is meaningful,
  • distribution of sequence lengths,
  • and differences by language and domain.

Vocabulary health

  • how many entries are extremely rare,
  • how balanced token frequencies are,
  • which domains dominate learned pieces,
  • and whether sensitive or accidental long strings became vocabulary entries.

Application behavior

  • retrieval quality under token-aware chunking,
  • code and JSON round trips,
  • chat-template correctness,
  • context overflow rate,
  • prompt cost,
  • and downstream model quality.

Reproducibility

Version together:

model checkpoint
tokenizer files
normalization and preprocessing config
special-token map
chat template
library version or compatibility test

A useful smoke-test corpus should include spaces, tabs, line breaks, combining marks, emoji, supported scripts, code, URLs, numbers, and all control tokens.


Runnable budgeting exercise

Start with counted tokens, not characters divided by four. The function below accepts counts from the exact tokenizer and chat template used by the target model. It is executable without downloading a model; replacing the example counts with actual encoded message lengths is the integration step.

def select_evidence(context_limit, fixed_input, output_reserve, margin, chunks):
    counts = (context_limit, fixed_input, output_reserve, margin)
    if any(type(value) is not int or value < 0 for value in counts):
        raise ValueError("limits and counts must be non-negative integers")
    if context_limit == 0:
        raise ValueError("context_limit must be positive")
    entries = list(chunks)
    seen = set()
    for chunk_id, token_count in entries:
        if not isinstance(chunk_id, str) or not chunk_id or chunk_id in seen:
            raise ValueError("chunk IDs must be non-empty unique strings")
        if type(token_count) is not int or token_count < 0:
            raise ValueError("token counts must be non-negative integers")
        seen.add(chunk_id)
    remaining = context_limit - fixed_input - output_reserve - margin
    if remaining < 0:
        raise ValueError("Fixed input and reserves exceed the context budget")
    chosen = []
    for chunk_id, token_count in entries:  # ranked by usefulness
        if token_count <= remaining:
            chosen.append(chunk_id)
            remaining -= token_count
    return chosen, remaining

chosen, unused = select_evidence(8192, 1500, 1500, 512,
                                [("policy", 2400), ("example", 1800),
                                 ("exception", 400)])
assert chosen == ["policy", "example", "exception"]
assert unused == 80

Here the evidence budget is 8192 − 1500 − 1500 − 512 = 4680; the selected passages consume 4,600 tokens. Increase the exception to 600 tokens and it no longer fits. If that exception is essential, replace a lower-value passage or narrow the task; dropping it silently could change the answer's meaning. This greedy example demonstrates accounting, not an optimal evidence-selection algorithm.

Count role markers, tool schemas, multimodal input under the provider's rules, and any model-specific output/reasoning budget. Check the input and output limits independently when the interface has separate caps. After packing, encode the final request again because joining pieces can change tokenization.

Worked interview: a multilingual request-budget gateway

Prompt: A RAG assistant serves English, Hindi and Japanese documents and supports two generation providers. Requests occasionally overflow context, and citations sometimes highlight the wrong characters. Design the tokenization and request-assembly layer. All volumes and prices below are interview assumptions.

1. Functional requirements

  1. Accept structured messages, authorized retrieved passages and optional supported media.
  2. Count against the selected provider/model and reserve output before sending generation requests.
  3. Preserve required instructions, the current question and complete tool-call/result groups.
  4. Select evidence under the budget while retaining document version and source offsets for citations.
  5. Report context overflow clearly when required content alone cannot fit; do not silently remove it.
  6. Support a provider/model change through a versioned adapter and an evaluated rollout.

2. Non-functional requirements

  1. Planning load: one million requests/month and a peak of 100 requests/second.
  2. Proposed local assembly target: p95 below 50 ms for bounded inputs. Track remote counting latency separately.
  3. No locally known over-budget request is sent. Track provider rejection rates because remote counts and evolving wrappers can differ.
  4. Preserve source text required for code, citations and multilingual output; test the offset coordinate conversion.
  5. Keep tenant data and cached request counts isolated; do not store raw prompts in metrics.
  6. Compare task quality and successful-request cost by language, not only a blended token average.

3. Simple design, then failure analysis

A first version estimates tokens as characters / 4, sums passage counts and trims the oldest strings. It is small, but its assumptions fail:

Failure Why it happens Repair Tradeoff
Non-English requests overflow English character ratios are used as enforcement Count with the correct model adapter More CPU work or a remote round trip
Joined text exceeds the sum of parts Separators and boundary merges change tokenization Count the fully serialized candidate request Repeated counting needs bounded work
A tool result loses its call History is cut by arbitrary strings Treat protocol-related messages as one group Less flexibility in packing
Citation highlights split an emoji Python code-point offsets reach a UTF-16 browser unchanged Version source text and convert offset units More mapping metadata and tests
A provider switch keeps old limits Cached counts and wrapper rules omit model identity Pin the complete adapter version Separate caches and compatibility checks
A required legal exception disappears Evidence is selected only by a scalar rank Mark required evidence groups and reject/narrow when they cannot fit Some requests require a smaller scope

4. Refined request path

Architecture / visual model
flowchart TD A[Structured request<br/>authenticated scope] --> B[Choose pinned adapter<br/>model, tokenizer, processor, template] B --> C[Preserve required groups<br/>instructions, question, tool exchanges] C --> D[Authorized retrieval<br/>source versions and offsets] D --> E[Structure-aware evidence packing<br/>count joined candidates] E --> F[Count final request<br/>input, context and output caps] F -->|fits and quota reserved| G[Submit generation request] F -->|required content cannot fit| H[Explain limit or ask for narrower scope] F -->|optional content removable| E G --> I[Incremental decoding<br/>citation offset conversion] G --> J[Reconcile reported usage<br/>release unused budget reservation] B <--> K[Scoped count cache<br/>exact request and adapter identity] J --> M[Metrics by language and model<br/>overflow, latency, cost, task quality]
Read diagram source
flowchart TD
    A[Structured request<br/>authenticated scope] --> B[Choose pinned adapter<br/>model, tokenizer, processor, template]
    B --> C[Preserve required groups<br/>instructions, question, tool exchanges]
    C --> D[Authorized retrieval<br/>source versions and offsets]
    D --> E[Structure-aware evidence packing<br/>count joined candidates]
    E --> F[Count final request<br/>input, context and output caps]
    F -->|fits and quota reserved| G[Submit generation request]
    F -->|required content cannot fit| H[Explain limit or ask for narrower scope]
    F -->|optional content removable| E
    G --> I[Incremental decoding<br/>citation offset conversion]
    G --> J[Reconcile reported usage<br/>release unused budget reservation]
    B <--> K[Scoped count cache<br/>exact request and adapter identity]
    J --> M[Metrics by language and model<br/>overflow, latency, cost, task quality]
  1. Separate tokenizer limits. The embedding model's chunk limit and the generator's context limit may use different tokenizers. Check each where it applies; do not reuse an embedding-token count as a generation-token count.
  2. Preserve structure before packing. Keep required groups intact. Split oversized optional documents at valid source boundaries and retain their source/version/offset mapping. The count-only exercise above is an accounting aid, not proof that selected evidence answers the question.
  3. Use a bounded refinement loop. Recount the assembled request. Remove or shrink an optional group only if something actually changes; stop after a fixed attempt bound. Required-only overflow returns an error instead of looping forever.
  4. Freeze what was counted. After the final count, submit the same model, messages, tools and media configuration. If any field changes, recount. Provider-managed conversation state must be included through its supported counting contract or handled conservatively.
  5. Keep caches exact and private. Scope a count-cache entry by authenticated tenant plus a hash of the complete count request and adapter version. Token IDs and hashes are not anonymization. Use a retention policy and avoid caching raw secrets in logs.
  6. Reserve money separately from context. The count endpoint does not reserve tokens or spend. Atomically reserve an upper-bound request budget against the tenant's remaining allowance. Reconcile reported usage; an uncertain request stays reserved until resolved under policy. Do not release its reservation merely because the client disconnected.
  7. Roll out and roll back together. Canary the adapter with the model route, counting rules and citation mapping. Test languages, code, combining marks, emoji, tool exchanges, media and exact boundaries. Keep the prior route available and monitor actual usage drift.

5. Work a context example

For a hypothetical 32,768-token combined context:

Reservation Tokens
Required instructions, tools, current question and required history 6,000
Output reserve 2,048
Measured wrapper/counting margin 512
Remaining evidence allowance 24,208

Four individually counted passages of 6,000 tokens total 24,000. That does not prove they fit: if the final assembled input adds 300 tokens of labels and separators, the evidence contribution becomes 24,300, exceeding the allowance by 92. Reduce an optional passage at a valid boundary and recount the full request. A different tokenizer can change every number in this table.

For local counting capacity, suppose measured mean CPU time is 20 ms per request, including all recounts. At 100 requests/second, demand is two CPU-seconds/second. At 60% target utilization, at least four CPU cores are needed for this stage, before failure headroom. A p95 target is not a mean service time; this calculation needs measured mean work. Remote count calls additionally need bounded concurrency, timeout handling and provider rate-limit capacity.

6. Compare full costs

Assume one million monthly requests, 500 billable output tokens/request and illustrative rates of $2/million input tokens and $10/million output tokens. Quality-tested evidence packing reduces average input from 5,000 to 4,000 tokens. These rates illustrate the calculation and are not quoted current provider prices.

Monthly item Current assembly Revised gateway
Input inference $10,000 $8,000
Output inference $5,000 $5,000
Gateway, queue, storage and monitoring $400 $700
Operations $1,200 $1,500
Quality review $1,800 $1,800
Build cost amortized over six months $0 $800
Effective total $18,400 $17,800

The new implementation costs $4,800 once (40 hours × $120). Recurring savings before that amortization are $1,400/month, so simple payback is about 3.43 months. During six-month amortization the effective saving is $600/month. Do not subtract the build cost both upfront and again when calculating cash payback.

The revised total is $17.80 per 1,000 submitted requests, or $18.35 per 1,000 successful requests at an assumed 97% outcome success rate. If shorter evidence reduces answer quality, the optimization may lose money despite a smaller input bill. Counting-service charges, media and extra tool/reasoning usage must be added when the chosen provider bills them; this example assumes a text-only workload without those extra charges.

7. Closing remarks

“I would version the complete model-input contract, preserve required message groups and source offsets, and enforce limits on the final serialized request. I would measure cost and quality by language and treat remote counting as a separate dependency. A smaller token count is useful only if the request remains correct, complete enough and cheaper per successful outcome.”

Interview tip: When asked for a token estimate, distinguish an approximate planning count, an enforcement count for the actual model input, and the provider's reported billable usage.

19. Interview Questions

1. Why do language models use subword tokens instead of words?

A word vocabulary becomes enormous and still cannot cover every name, typo, or new word. A sufficiently complete character alphabet or byte fallback covers rare text but creates longer sequences. Subwords keep common text compact while spelling rare text from smaller reusable units.

2. Compare BPE, WordPiece, and Unigram

BPE begins with small units and learns a sequence of pair merges. WordPiece uses a likelihood-oriented training score and commonly applies greedy longest-match encoding. Unigram begins with many candidate pieces, assigns probabilities, and prunes the vocabulary while choosing likely segmentations. SentencePiece is a toolkit that can implement BPE or Unigram.

3. Why can two models charge different tokens for the same text?

They may use different normalization, pre-tokenization, vocabularies, algorithms, byte policies, and chat wrappers. Token count belongs to a particular serialized request and tokenizer version, not to the visible sentence alone.

4. Why can a larger vocabulary reduce one cost and increase another?

Larger vocabularies often shorten sequences, reducing attention positions. They enlarge embedding and output tables and can increase final projection work. The useful balance depends on the training corpus and workload.

5. How would you chunk documents for RAG?

Parse semantic units first, count them with the production tokenizer, pack them under a measured budget, split oversized units at smaller valid boundaries, preserve offsets and source metadata, and tune overlap through retrieval evaluation.

6. How would you evaluate multilingual fairness in tokenization?

Compare coverage and fertility by language and script on representative traffic, then measure downstream quality and cost. Report tokens per byte, character, grapheme, or word with a clearly defined denominator. Do not treat one English-centric average as universal.

7. What breaks if a tokenizer vocabulary is reordered?

Reassigned IDs select different embedding rows and output logits. Unless the corresponding weights and ID-dependent metadata are remapped consistently, the model-tokenizer contract is corrupted.

8. Why is exact character counting awkward for an LLM?

The model usually receives subword or byte-derived token positions rather than a guaranteed list of user-visible grapheme clusters. It can infer spelling patterns, but exact counting is better handled by a deterministic text tool.

9. Can decoding at a token boundary corrupt Unicode?

Yes. A byte-based token can hold only part of a multi-byte UTF-8 character. Preserve source text or carry decoder state across pieces. A valid character boundary can still split a multi-code-point grapheme.

10. Does SentencePiece guarantee the original whitespace returns unchanged?

No. Its space marker represents normalized text. The configured normalizer may collapse repeated spaces or alter Unicode forms before segmentation. Test round trips against the preservation requirement.

11. Why not sum independently counted messages to enforce context limits?

Wrappers, separators, tool schemas and boundary-dependent segmentation can change the total. Count the complete request that will actually be sent, reserve output, and verify any separate input/output caps.

12. Can two tokenizers with the same vocabulary size be swapped?

No. Their ID mappings, segmentation, normalization and special-token rules can differ. Match the exact checkpoint contract; equal table dimensions do not imply compatible meanings.

13. A tokenizer emits fewer tokens in Hindi. Is that enough to choose its model?

No. Measure task quality, coverage, context retention, latency and full cost. Tokenizer fertility needs a defined word segmentation; cross-language efficiency comparisons should also use clearly labeled byte/code-point ratios or aligned tasks.

14. Does the token-count endpoint prove a cache discount or reserve quota?

No. Counting estimates input size. Actual cache hits, output usage, rate limits and billing follow the generation contract. Reserve application spend atomically and reconcile the final reported usage.

15. Does spelling a word with spaces force one token per letter?

No. It changes the text and may help some models, but the tokenizer still chooses its own pieces. Use a deterministic code-point or grapheme-counting tool with a clear definition of “character.”


20. Compact Reference

Mental model

raw text
  → normalize
  → establish candidate boundaries
  → select learned pieces
  → map pieces to IDs
  → add structural IDs and metadata
  → look up model embeddings

Algorithm map

Name Remember this sentence
BPE Repeatedly learn useful adjacent merges from small base units.
Byte-level BPE Apply the BPE idea with complete byte coverage.
WordPiece Learn useful word pieces and commonly encode with greedy longest matching.
Unigram Score alternative segmentations with piece probabilities and prune a large candidate vocabulary.
SentencePiece Train BPE or Unigram directly from raw strings with explicit normalized-space markers.

Final notes

  1. Define the unit: token, token ID, byte, code point and grapheme are different.
  2. Preserve the contract: model weights, tokenizer, processor, template and normalization travel together.
  3. Count the actual request: joined input, wrappers, tools, media and output reserves all matter.
  4. Preserve the document: structure and source offsets matter more than arbitrary token slices.
  5. Measure outcomes: fewer tokens do not prove better quality or lower total operating cost.

System-design checklist

  • Exact tokenizer version matches the checkpoint.
  • Normalization and special-token rules are versioned.
  • Chat/tool templates are applied once.
  • Complete serialized requests are counted.
  • Output space is reserved before filling the context.
  • RAG chunks respect structure and token budgets.
  • Offset semantics are tested.
  • Multilingual, code, and noisy inputs are measured separately.
  • Streaming decode handles partial byte sequences.
  • Tokenizer changes trigger compatibility and quality tests.

21. Engineering References

These references connect the worked examples to tokenizer APIs and the papers that describe the algorithms.

Maintained explainers and tools

  1. Hugging Face. Tokenization algorithms — current explanations of BPE, byte-level BPE, Unigram, SentencePiece, WordPiece, word-level, and character-level tokenization. Read the reference

  2. Hugging Face Tokenizers. The tokenization pipeline — normalization, pre-tokenization, model, post-processing, and decoding. Read the reference

  3. OpenAI Developer Cookbook. How to count tokens with tiktoken — executable counting examples and model-encoding guidance. Read the reference

  4. OpenAI. Tokenizer tool — an interactive way to inspect one provider's token boundaries. Read the reference

  5. Google. SentencePiece repository — maintained implementation notes and training examples. Read the reference

Primary papers

  1. Sennrich, R., Haddow, B., and Birch, A. Neural Machine Translation of Rare Words with Subword Units (2016). Read the reference

  2. Kudo, T. Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates (2018). Read the reference

  3. Kudo, T. and Richardson, J. SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing (2018). Read the reference

  4. Schuster, M. and Nakajima, K. Japanese and Korean Voice Search (2012), an early published WordPiece reference. Read the reference


Previous: LLM Fundamentals | Next: Attention Mechanisms

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← LLM Internals: How a Model Learns and Answers
NEXT LESSONAttention Mechanisms: How Tokens Share Information →

Explore the diagram