One visible word can become several units.
Tokenizers normalize and segment text, apply a tokenization method, add any required special tokens, and return model inputs such as token IDs. The result depends on the tokenizer’s vocabulary and rules. Words, punctuation, whitespace, spelling and language are not interchangeable with token boundaries.