A tokenizer running entirely in your browser. Count tokens and compare how different large language model vocabularies split the same text.
Above: the same five pieces in three vocabularies, with a different ID every time.
Below: loads tokenizer.json and tokenizer_config.json from any repository on
Hugging Face. Also handy for debugging prompt templates.
#, so even long texts make a working link?text=your%20text&models=model1,model2,model3 format" ." keeps its space. Tick "Clean up spaces before punctuation" on a card to see it the way the model's decoder would tidy it (".", "don't"), next to that model's own default. Only the text changes, never the IDs or the count<ruby> elements with text above and token numbers belowCompressionStream('deflate-raw') into #text=..., which browsers never send to the server, so links don't run into its URL length limit. ?text=...&models=... also works