What is a token in AI? The 4 kinds and what each one costs

A token in AI is a chunk of text, about three quarters of a word, and it's the unit AI models read, limit and charge in. Here's how tokens work, the four kinds you pay for and how to use fewer.

Navid Moazzezby Navid Moazzez·Updated Oct 2, 2026·7 min read·

A token in AI is a chunk of text, about three quarters of a word, and it's the unit AI models read, remember and charge you in.

You never see tokens in ChatGPT or Claude. But they decide how much fits in one chat, when your plan runs out and what every answer costs.

I build and run my business with Claude Code every day. Here's my real token usage, live:

5,778 contributions on GitHub in the last year

In this guide, I'll show you:

  • What a token is, and why AI reads tokens instead of words
  • How to turn words into tokens, and tokens back into words
  • The four kinds of tokens you pay for
  • What tokens cost, and why output costs five times as much as input
  • How to count tokens before you send anything
  • How to use fewer tokens without getting worse answers

Once you see your work in tokens, the limits and the bills stop being a surprise.

key_takeaways.mdTL;DR

Key takeaways

A token is a piece of text, about 4 characters or three quarters of a word in English. 100 tokens is about 75 words.
AI models read your text, set their limits and charge for API use in tokens, from your message and your files to the answer they write.
There are four kinds: input tokens, output tokens, cached tokens and reasoning tokens.
On Claude, output tokens cost five times as much as input tokens, and cached input costs a tenth of the normal price or less.
A long chat costs more with every question, because every new message sends the whole conversation again.

What is a token?

A token is the smallest piece of text an AI model works with. It can be a whole word, part of a word, a single character or a punctuation mark.

The model never reads your words. It splits your text into tokens, works on those, and writes its answer back one token at a time.

For example, a short, common word like "the" is usually one token. A long or rare word breaks into several pieces. And code breaks into many, because every bracket and underscore counts.

Even a space changes things. In OpenAI's own example, "red", "Red" and " red" with a space in front can each be a different token.

Why does AI use tokens instead of words?

Each model has its own encoding, a fixed list of tokens it knows, each with its own ID number. Common words get a token of their own, and everything else is built from smaller pieces.

That's how a model can read any text at all: new words, names, typos, other languages and code. If it only knew whole words, every word it had never seen would be a dead end.

It's also why two models can count the same sentence differently, because each one uses its own list.

How many words is a token?

In English, one token is about 4 characters, or three quarters of a word. OpenAI and Anthropic both give that rule of thumb.

So 100 tokens is about 75 words, and 1,000 tokens is about 750 words:

TokensAbout this many English words
10075
1,000750
10,0007,500
100,00075,000
1,000,000750,000

It works the other way too. Multiply words by 1.33 to get tokens, so a 500-word prompt is about 670 tokens, not 500.

Files add up fast. Anthropic's own estimates put an average web page at about 2,500 tokens and a research paper PDF at about 125,000.

Three things change the count:

  • Language – The 4-characters rule is for English. Other languages have their own ratio, so the same message can use more or fewer tokens.
  • Code – Symbols, brackets and indentation each take tokens, so a line of code uses more tokens than a line of prose.
  • The model – Each model family splits text its own way. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text than earlier Claude models.

The four kinds of tokens

When you use AI through an app or an API, your tokens fall into four kinds, and each one is counted and priced on its own.

Input tokens

Input tokens are everything you send: your message, your files, your rules and the whole conversation so far.

That last part is the one people miss. Every new message sends the full history again, so your 40th message resends the first 39 with it.

That's why a chat that has run all afternoon costs more per question than the same question asked in a fresh one.

Output tokens

Output tokens are what the model writes back, and they're the expensive half. On Claude, they cost five times as much as input tokens.

So asking for a summary instead of a full rewrite, or for just the lines that change instead of the whole file, saves more than trimming your prompt.

Cached tokens

Cached tokens are input the provider has already processed and can reuse, as long as the start of your request hasn't changed.

Your rules and context files are the same at the start of every message, so they get stored and charged at a fraction of the price. On Claude, reading from the cache costs a tenth of the normal input price or less, and the cache lasts 5 minutes, renewed each time it's used.

The catch is that the cache only matches from the start. Change one character near the top, like a date or a timestamp, and everything after it is charged in full again.

Reasoning tokens

Reasoning tokens are the thinking a reasoning model does before it answers. You don't see them, but they count, and they're billed as output tokens.

So a three-line answer can still use a lot of tokens if the model thought hard to get there.

What do tokens cost?

API prices are per million tokens, and input and output each have their own price. These are Claude's, from Anthropic's pricing page on October 1, 2026:

Claude Haiku 4.5
Input
$1
Output
$5
Cached input
$0.10
Claude Sonnet 5
Input
$2
Output
$10
Cached input
$0.20
Claude Opus 5.5
Input
$4
Output
$20
Cached input
$0.20

A single answer is cheap. A 1,000-word article is about 1,330 output tokens, which costs about 1.3 cents on Claude Sonnet 5.

The cost comes from repetition. Say an agent has 100,000 tokens of files and history in its context, and you ask it 30 more questions. That's 3 million input tokens, or $6 on Sonnet 5. With caching, most of those are read at a tenth of the price, and the same questions cost under $1.

On a flat plan like Claude Pro or Max, you don't pay per token, but tokens still drive your usage limit. Anthropic lists message length, file size, conversation length, model choice and effort level among the things that make you reach it faster.

I'm on Claude Max, and my live usage card above shows what the same tokens would have cost on the API.

How to count tokens

You don't have to guess. Here's how to see the real number:

  • OpenAI's Tokenizer – Paste any text into OpenAI's Tokenizer to see how it splits and how many tokens it is.
  • Claude's token counting – Anthropic's API has a token counting endpoint that counts a full request, files and tools included, before you send it.
  • Claude Code – Type /context to see what's filling your context window, from your rules files to the conversation itself.
  • The usage report – After every request, the API returns the exact number of input and output tokens it used.

How to use fewer tokens

Most of the savings come from a few habits:

Use fewer tokens0/6

Common token mistakes

These are the token mistakes I see most:

  • You count words as tokens. A 500-word prompt is about 670 tokens.
  • You keep one chat open all day. Every new message resends the whole conversation.
  • You ask for a whole file back when you only need one change. Output costs five times as much as input.
  • You put a date or anything else that changes at the top of a prompt. That breaks the cache for everything after it.
  • You compare models only on their price per million tokens. Models split the same text differently, so compare what a real task costs.

FAQs about tokens in AI

Here are the questions people ask most about AI tokens, with a direct answer to each.

A token is a chunk of text, about 4 characters or three quarters of a word in English.

AI models read your text, set their limits and charge for API use in tokens, from your message to the answer they write.

In English, 1,000 tokens is about 750 words.

It works the other way too, so a 500-word prompt is about 670 tokens.

Writing text takes a model more work than reading it, so output is priced higher.

On Claude, output tokens cost five times as much as input tokens.

Cached tokens are input the provider has already processed and can reuse while the start of your request stays the same.

On Claude, reading from the cache costs a tenth of the normal input price or less.

Reasoning tokens are the thinking a reasoning model does before it answers.

You don't see them, but they're billed as output tokens.

Paste your text into OpenAI's Tokenizer to see how it splits and how many tokens it is.

In Claude Code, type /context to see what's filling your context window.

No, each model family splits text its own way.

Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text than earlier Claude models.

Final thoughts on tokens

Tokens are the number AI tools run on, even when the app hides it from you.

You don't need to count every one. Remember the rule of thumb, about three quarters of a word, and remember where tokens go: long chats, big files and long answers.

Keep reading:

Next time a chat starts forgetting things or your plan runs out early, check the tokens first.

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

Related free tools

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers