AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→

Q: What are tokens?

Tokens are the units of text an LLM actually reads and writes. Not letters, not words. Tokens.

Before your text reaches the model, a tokenizer splits it into pieces. Common short words become one token each. Longer or rarer words get split into several. For example, "cat" is one token, while "unbelievable" gets split into pieces like "un", "believ", "able". A useful rule of thumb for English: one token is about 4 characters, and 1,000 tokens is roughly 750 words.

Each model family has its own tokenizer. You can see this for yourself with OpenAI's online tokenizer: paste any text and it shows you exactly how it splits into tokens, with a live count.

Why should you care? Because everything is measured in tokens. API pricing is per token, with input and output priced differently. The context window (how much the model can hold at once) is a token limit. Prompt caching discounts are per token. When your app gets expensive or hits limits, tokens are almost always where you look first.

← Back to the full FAQ