Subword Tokenization Conceals Character Boundaries: Why Autoregressive Models Fail at Letter Counting, Spelling, and String Reversal
Large language models struggle with simple character-level tasks because their input mechanism compresses text into opaque subword chunks. By prioritizing computational efficiency over character visibility, tokenizers structurally hide the internal letters of words before the neural network even begins processing them.
By Sofia Matos
In short
- Large language models fail at spelling and letter counting because subword tokenization compresses text into opaque integer IDs, hiding internal character boundaries.
- Byte Pair Encoding (BPE) is used to keep sequence lengths short and compute costs low, but it sacrifices granular character visibility to achieve that efficiency.
- Tokenization heavily favors English; non-English languages often fragment into more tokens, increasing API costs and reducing usable context windows.
Why do autoregressive models fail at counting the Rs in "strawberry" or reversing a simple string? The answer lies entirely in how they read text. Large language models do not see letters; they see opaque integer IDs called tokens.[2]
Before a prompt ever reaches the neural network's attention layers, a separate software module called a tokenizer chops the text into chunks. This preprocessing step structurally hides the internal character boundaries of words.[5]
When a user asks a model to count letters, they are asking it to reason about internal structure that its input representation deliberately discarded. The failure is not a lack of reasoning, but a literal blind spot in the data.[4]
The Mechanics of Subword Tokenization
The dominant approach in modern natural language processing is subword tokenization, specifically Byte Pair Encoding (BPE). Originally developed for data compression in the 1990s, BPE balances vocabulary size with sequence length.[1]
BPE works by iteratively merging the most frequent adjacent byte pairs in a training corpus into single, new symbols. Common words become single tokens, while rare words are split into smaller pieces.[1]
For example, the word "strawberry" does not enter the model as ten separate letters. Instead, a standard BPE tokenizer might split it into two tokens: "straw" and "berry".[2]
The model's embedding layer assigns a single high-dimensional vector to the token "straw". It has no mechanism to look inside that vector and extract the individual characters s-t-r-a-w.[2]
Once the text is tokenized, the original character boundaries are thrown away. From that point forward, the transformer architecture operates entirely on vectors and matrices representing those subword chunks.[5]
The Trade-Off Between Compute and Visibility
If subword tokenization causes these blind spots, why not simply use character-level tokenization? The answer comes down to computational efficiency and the strict limits of transformer architecture.[1]
Treating every character as an individual token would solve the spelling problem, but it would cause sequence lengths to explode. A 500-word document would suddenly require processing around 2,500 individual tokens.[1]
Because the computational cost of a transformer's attention mechanism scales quadratically with sequence length, character-level tokenization is prohibitively expensive. It would drastically shrink the usable context window of any model.[1]
Word-level tokenization sits at the other extreme, but it fails whenever it encounters an out-of-vocabulary word. It also struggles immensely with morphologically rich languages or languages without clear spaces, like Chinese or Japanese.[1]
Subword tokenization acts as the necessary compromise. It keeps the vocabulary size manageable and the sequence length short, but it sacrifices granular character visibility to achieve that efficiency.[1]
Arithmetic and the Number Splitting Fix
This tokenization blind spot extends beyond spelling and string reversal. Historically, it has been the primary reason language models struggled with simple multi-digit arithmetic.[3]
If a tokenizer merges numbers based on frequency, a number like 127 might become a single token, while 677 might split into two separate tokens. The model is forced to learn addition rules for arbitrary chunks of digits.[1]
Research on frontier models demonstrates that arithmetic performance is highly dependent on how numbers are tokenized. When numbers are split inconsistently, the model's computations become approximate rather than systematic.[3]
To fix this, many modern tokenizers now include explicit rules to split multi-digit numbers into individual digit tokens. This prevents BPE from merging them into arbitrary chunks based on training corpus frequency.[4]
By forcing the tokenizer to represent 1024 as four separate tokens, the model can finally align the digits properly in its attention layers. This simple preprocessing change dramatically improves arithmetic accuracy.[4]
The Illusion of Spelling Memorization
A common paradox arises when testing these models: if they cannot see internal characters, how can they accurately spell out a word when asked? Models can often spell words with over 94% accuracy.[5]
The answer is memorization during pretraining, not dynamic extraction. The model has seen the sequence of letters associated with the word countless times in its training data, allowing it to predict the spelling sequence.[5]
However, probing analyses reveal that the token embedding layer itself does not encode complete character-level information beyond the first letter. The model is reciting a memorized fact, not looking at the token's internal structure.[5]
This is why a model can easily generate a word's spelling on demand, but will confidently fail if asked to identify the fifth letter of that same word. It cannot reliably extract positional character data that was never embedded.[5]
Multilingual Disparities and Algorithmic Bias
The consequences of subword tokenization are not distributed equally. Because BPE is a frequency heuristic, it merges whatever byte pairs happen to co-occur most often in the training corpus.[4]
Since most training corpora are overwhelmingly English, the resulting token vocabularies are heavily optimized for English text. Common English words get compressed into single tokens efficiently.[4]
For non-English languages, especially those with different scripts or complex morphology, the tokenizer is far less efficient. A single word might shatter into half a dozen fragmented byte tokens.[4]
This creates a measurable disparity. The same semantic sentence translated into different languages can result in tokenized sequences that differ in length by a factor of up to fifteen.[4]
Users working in underrepresented languages effectively pay more per query under token-based pricing models. They also hit context window limits much faster, all due to a vocabulary decision made before the model was even trained.[4]
The Path Forward for Model Architecture
Tokenization remains the invisible layer between human language and machine learning. It is a fundamental design choice that dictates exactly what the neural network is permitted to process.[1]
As models move toward more complex reasoning tasks, the limitations of BPE are becoming a bottleneck. The rigid reliance on a static, pre-computed vocabulary prevents true end-to-end optimization of the language model.[5]
Researchers are actively exploring alternative architectures, such as dual-tokenization models or byte-level models that bypass BPE entirely. However, the computational cost of these approaches remains a significant hurdle.[5]
Until the underlying architecture evolves to handle longer sequences efficiently, subword tokenization will remain the standard. When an autoregressive model fails at a simple character task, it is operating exactly as designed.[2]
How we did this
- Method
- Comparing the sequence length and internal boundary retention of the word 'strawberry' across character-level and subword-level tokenization schemes to quantify structural data loss.
- What we found
- Subword tokenization achieves an 80% sequence compression rate on common words but structurally deletes 88% of the internal character boundaries from the input matrix, proving that letter-counting failures are a data-loss issue before the neural network begins processing.
- What we worked from
- Character-level sequence length for 'strawberry': 10 tokens (9 internal boundaries) — Tom Archer Blog
- BPE subword sequence length for 'strawberry': 2 tokens (1 internal boundary) — Async Thinking
- Limits of this analysis
- This analysis uses a single English word as a proxy; compression rates and boundary loss vary significantly across different languages, rare words, and specific tokenizer vocabularies.
Key terms
- Tokenization
- The preprocessing step that converts raw text into a sequence of discrete integer IDs that a language model can process.
- Byte Pair Encoding (BPE)
- A subword tokenization algorithm that iteratively merges the most frequent adjacent byte pairs in a training corpus into single tokens.
- Out-of-Vocabulary (OOV)
- A scenario where a model encounters a word it was not explicitly trained to recognize, causing word-level tokenizers to fail.
- Embedding Layer
- The initial neural network layer that maps each discrete token ID to a high-dimensional vector representing its semantic meaning.
Frequently asked
Why don't models just use character-level tokenization?
Processing every character individually causes sequence lengths to explode, which drastically increases the computational cost of the transformer's attention mechanism. It would severely limit the usable context window of any model.
How do newer models fix the arithmetic problem?
Many modern tokenizers now include explicit rules to split multi-digit numbers into individual digit tokens. This prevents the tokenizer from merging numbers into arbitrary chunks, allowing the model to align digits properly.
Does tokenization affect API costs?
Yes. Because commercial LLM APIs charge per token, users working in languages that tokenize inefficiently end up paying significantly more for the same semantic content compared to English users.
Viewpoints in depth
AI Architects
Prioritize computational efficiency and expanded context windows through aggressive subword compression.
For the engineers designing frontier models, subword tokenization is a necessary compromise to keep computational costs manageable. Because the attention mechanism in a transformer scales quadratically with sequence length, processing text character-by-character would drastically shrink the usable context window. By compressing common words into single tokens, architects can feed the model significantly more information per query, accepting minor character-level blind spots as a worthwhile trade-off for broader semantic reasoning capabilities.
NLP Linguists
Argue that frequency-based tokenization destroys morphological structure and systematically disadvantages non-English languages.
Linguists point out that Byte Pair Encoding is a purely statistical heuristic with no understanding of grammar or morphology. It merges whatever byte pairs happen to co-occur most often in the training corpus, which is overwhelmingly English. As a result, the vocabulary reflects statistical accidents rather than linguistic structure. This creates a systemic disadvantage for morphologically rich languages, where words shatter into inefficient fragments, driving up API costs and degrading model performance for non-English speakers.
AI Safety Researchers
Highlight how tokenization artifacts create unpredictable edge cases and adversarial vulnerabilities.
Safety researchers view tokenization as a critical vulnerability layer. Because the tokenization process is disjoint from the model's actual training, it introduces rigid artifacts that the neural network cannot override. Unusual character combinations, trailing whitespace, or specific glitch tokens can bypass safety filters or cause erratic behavior simply because the tokenizer parses them into unexpected integer sequences. They argue that true end-to-end optimization is impossible as long as this opaque preprocessing step remains in place.
- AI Architects
- Prioritize computational efficiency and expanded context windows through aggressive subword compression.
- NLP Linguists
- Argue that frequency-based tokenization destroys morphological structure and systematically disadvantages non-English languages.
- AI Safety Researchers
- Highlight how tokenization artifacts create unpredictable edge cases and adversarial vulnerabilities.
Perspectives this story doesn't cover
- Hardware engineers designing next-generation chips to handle longer sequence lengths
- End-users paying higher API costs for non-English queries
Sources
[1]Tom Archer BlogAI ArchitectsSubword Tokenization: The Goldilocks Solution
Read on Tom Archer Blog →
[2]Async ThinkingAI Safety ResearchersTokens, Not Letters
Read on Async Thinking →
[3]arXivTokenization counts: the impact of tokenization on arithmetic in frontier LLMs
Read on arXiv →
[4]My Written WordNLP LinguistsWhy Tokenization Explains LLM Failures
Read on My Written Word →
[5]Factlen Editorial TeamSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
More in Artificial Intelligence
See all →Model Architecture
The Mechanics of Mixture of Experts: How LLMs Route Queries to Specialized Sub-Networks
5 sources
Model Architecture
Startup Inception Unveils Diffusion-Based LLM, Claiming Several-Fold Speed and Halved Cost Over Conventional Models
4 sources
Biosecurity
Google DeepMind Unveils SynthID Bio to Watermark AI-Designed Proteins for Biosecurity Screening
5 sources
Frontier Models
Google Releases Gemini 4 Argon Frontier AI Model With 1-Million-Token Output Limit
7 sources
Comments
Every angle. Every day.
Get Artificial Intelligence stories with full source coverage and perspective breakdowns, free every day.




