A token is the unit a language model reads and writes. Before a model sees your prompt, a tokenizer cuts the text into pieces and replaces each piece with a number from a fixed vocabulary. Providers bill by these pieces, rate limits count them, and a model’s context window is a maximum number of them. This article shows what the pieces look like for real text, why the same sentence gives different counts on different models, and how to count tokens yourself. Every count below comes from the ZeroTool AI Token Counter and matches the providers’ own tokenizers.

A token is a piece of text with a number

Here is the question of this article, split by OpenAI’s o200k_base, the tokenizer of GPT-4o, GPT-4.1 and the GPT-5 family:

PieceWhat is a token in AI?
Token id482738226166023062083730

That is 7 tokens for 22 characters. Common English words are single tokens, and the space before a word is part of the token. The ids are only positions in the vocabulary: the model never sees letters, only these numbers.

Pieces do not have to be whole characters. A tokenizer works on UTF-8 bytes, so a character the vocabulary does not hold is split into byte pieces. The parrot emoji 🦜 is 4 bytes in UTF-8 and becomes 3 tokens in o200k_base: F0 9F, A6 and 9C. None of them is a character on its own. The token view in the counter shows such pieces in hex inside ⟨⟩ boxes.

How a tokenizer decides where to cut

The tokenizers used by GPT, DeepSeek and most open models are byte-level BPE (byte pair encoding, from Sennrich et al., 2016). They work in two steps.

  1. Pre-split with a regular expression. The text is first cut into rough chunks: words with their leading space, runs of digits, punctuation, runs of white space. Merges never cross these chunk borders. o200k_base keeps up to three digits together; DeepSeek V4 splits digits into groups of one to three and also cuts runs of Chinese characters and Japanese kana into their own chunks.
  2. Merge byte pairs by rank. Each chunk starts as single bytes. The tokenizer repeatedly joins the adjacent pair that ranks highest in its learned list of merges, until no listed pair is left. The training data decided that list: frequent pairs got high ranks.

The vocabularies differ in size. cl100k_base has 100,277 entries, o200k_base 200,019 and DeepSeek V4 129,280, counting special tokens. A bigger vocabulary can hold more whole words from more languages, which is why the move from cl100k_base (GPT-4) to o200k_base (GPT-4o) cut the count for Chinese, Japanese and Korean text.

The same sentence, different counts

Counts are only comparable within one tokenizer. The same short texts on three of them:

TextCharacterso200k_basecl100k_baseDeepSeek V4
What is a token in AI?22777
你好,世界!6474
東京の天気は晴れ87107
什么是 AI 里的 token?167106

English barely changes between tokenizers; Chinese and Japanese change a lot. In cl100k_base the character 世 is not a token, so 你好,世界! spends two byte tokens on it.

The gap grows with length. The counter’s built-in Chinese sample, a 192-character code review request, is 111 tokens in DeepSeek V4, 133 in o200k_base and 172 in cl100k_base:

That is 111, 133 and 172 tokens for one text, a 55% spread. The providers’ own rules of thumb hide this. Anthropic’s pricing FAQ says a token is about 4 characters or 0.75 words in English; Google’s token guide says about 4 characters for Gemini; DeepSeek’s documentation says 1 English character is about 0.3 tokens and 1 Chinese character about 0.6. For the 192-character sample, “4 characters per token” predicts 48 tokens, less than half of any real count. Each provider adds that the actual count is what the model returns.

Claude and Gemini are a special case: their tokenizers are not published for local use. Anthropic states that Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text than earlier Claude models. So even within one provider, a count from last year’s model is not this year’s count.

Tokens set the price and the limits

API prices are per million tokens, separately for input and output:

cost = input_tokens × input_price / 1,000,000 + output_tokens × output_price / 1,000,000

With prices checked on 2026-10-01, the 192-character Chinese sample plus 1,000 output tokens costs $0.030665 on gpt-5.5 ($5 input and $30 output per million, 133 tokens) and $0.0012333 on deepseek-flash at peak rates ($0.30 and $1.20, 111 tokens). Output dominates both: for short prompts, the length of the answer decides the bill.

Two rules change the arithmetic for long prompts. OpenAI prices gpt-5.5 and GPT-6 requests above 272,000 input tokens at 2× input and 1.5× output for the whole request, and Google charges more above 200,000 input tokens on gemini-3.1-pro-preview. And every model has an input limit: 128,000 tokens for gpt-4o, 272,000 for gpt-5.4-mini, 8,192 per input for text-embedding-3-small. When you chunk documents for embeddings, count the chunks with cl100k_base, which OpenAI’s embeddings guide names for the text-embedding-3 models.

Counting tokens from Python and the provider APIs

For OpenAI models, the reference implementation is tiktoken. Version 0.14.0:

import tiktoken

enc = tiktoken.get_encoding("o200k_base")
ids = enc.encode_ordinary("What is a token in AI?")
print(len(ids), ids)
# 7 [4827, 382, 261, 6602, 306, 20837, 30]

print(tiktoken.encoding_name_for_model("gpt-5.5"))  # o200k_base
tiktoken.encoding_name_for_model("gpt-6-sol")       # KeyError: Could not automatically map gpt-6-sol ...

encoding_for_model() matches name prefixes, so any gpt-5… name maps to o200k_base. GPT-6 has no entry: OpenAI has not published its tokenizer, and a tool that shows an exact GPT-6 count is guessing. Use encode_ordinary() for user text. Plain encode() raises ValueError when the text contains <|endoftext|>, which happens as soon as someone pastes tokenizer documentation into a prompt.

For DeepSeek, the token usage page links an offline package with the official tokenizer.json. Load it with the tokenizers library:

from tokenizers import Tokenizer

tok = Tokenizer.from_file("deepseek_v4_tokenizer/tokenizer.json")
enc = tok.encode("什么是 AI 里的 token?")
print(len(enc.ids), enc.ids)
# 6 [25043, 7703, 223, 7210, 17840, 1148]

The package’s own demo script uses transformers. On 2026-10-01, transformers 5.18.0 returned an empty list for Chinese text with this package, while transformers 4.57.6 and tokenizers 0.23.2 returned the ids above. Pin transformers 4.x or use tokenizers directly.

Claude and Gemini need their APIs. Anthropic’s POST /v1/messages/count_tokens takes the same body as a message request and is free within its own rate limits; its documentation calls the result an estimate. Gemini’s models.countTokens does the same for Google. OpenAI also has POST /v1/responses/input_tokens, and it is the only way to include what local tokenizers cannot see: OpenAI’s token counting guide says the count includes formatting tokens for message roles and boundaries, and covers images, files and tool definitions.

Where local counts go wrong

  • Message overhead. A local count measures text. A chat request adds role and separator tokens, a system prompt and tool schemas. Compare a local count with the API’s usage only for the same bare text.
  • JavaScript ports. js-tiktoken reads the regex shorthand \s with JavaScript’s rules, which include U+FEFF and exclude U+0085; tiktoken uses Unicode White_Space. Text that starts with a byte-order mark can split differently. Its merge loop is also quadratic: in our test, 20,000 Chinese characters with no punctuation took 229 seconds in Node.js. The ZeroTool counter rewrites \s and merges with a priority queue, and its test suite compares ids with tiktoken and tokenizers on more than 1,500 strings.
  • Fixed ratios. A count from “characters ÷ 4” was off by more than half for the Chinese sample above. Some online counters use such a ratio for every non-OpenAI model; on 2026-10-01 one showed 14 tokens for a sentence that is 35 tokens in o200k_base and 26 in DeepSeek V4.

Further reading