AI Token Counter

Characters, words and an estimated token count for prompts and documents, with the method shown.

Free tools that run in your browser. Nothing you type, open or create is uploaded to NerdBible or anyone else.

How this estimate was worked out
Part of the textFoundRule usedTokens

These are estimates. Every AI model family splits text with its own tokenizer, and the same text can come out as a noticeably different number of tokens in each one; newer tokenizers are generally more efficient with languages other than English. For an exact count use the tokenizer or token counting tool published by the model's provider. Your text stays in this browser and is not sent anywhere.

How to use it

  1. Paste text, a prompt or code into the box, or open a text file from your device.
  2. Read the characters, words, lines and the estimated token count as you type.
  3. Open "How this estimate was worked out" to see how each part of the text was counted.

Questions

What is a token?

Language models do not read letters or whole words. Their tokenizer cuts text into pieces called tokens: a common English word is often one token, a long or rare word several, and punctuation, numbers and spaces in code get their own. Context limits and usage are measured in tokens.

Why is this an estimate and not an exact count?

An exact count needs the exact tokenizer of the model you use, and each provider has its own, sometimes more than one. This page uses a published rule of thumb instead, so it works for any text without downloading large vocabulary files. Expect the real figure to land within the range shown for ordinary English, and to vary more for other scripts.

What rule of thumb is used?

The common guide for English is about four characters, or three quarters of a word, per token. This page refines that by counting each kind of text separately: short words as one token, long words as more, numbers in groups of three digits, punctuation runs, line breaks, roughly one token per Chinese, Japanese or Korean character and fewer characters per token for scripts such as Cyrillic, Arabic or Devanagari. Dense code gets a reduction because tokenizers merge common symbol patterns.

Are characters counted the way models count them?

Characters here are what you would see: an emoji or an accented letter counts once, even when it is stored as several code units. Without spaces leaves out spaces, tabs and line breaks.

Is my text uploaded?

No. Counting happens in your browser as you type, and nothing you paste or open is sent to NerdBible or to any AI provider.

About this tool

About the AI token counter

Language models read text as tokens, not words, and their context limits and usage are measured in tokens. Paste a prompt, a document or some code here to see its characters, words and lines, and an estimate of its token count with a likely range.

Why it is an estimate

Each model family cuts text up with its own tokenizer, and the same text can produce noticeably different counts in each. An exact count needs that exact tokenizer, which the provider of each model publishes. This page uses a documented rule of thumb that works for any text, so treat the figure as a guide for planning, not as a billing number.

How the estimate is made

  • The common guide for English is about four characters, or three quarters of a word, per token.
  • Words, numbers, punctuation, line breaks and spaces are counted separately, each with its own rule.
  • Chinese, Japanese and Korean characters count roughly one token each; Cyrillic, Arabic, Devanagari and other scripts get fewer characters per token.
  • Text dense with symbols, such as code, gets a reduction, because tokenizers merge common symbol patterns.

The breakdown table shows every rule applied to your text. Nothing you paste is sent anywhere.

Related tools