Models like ChatGPT struggle with basic string manipulation because their fundamental unit of understanding is a "token," which can be a whole word or a sub-word. If a word like "Strawberry" is tokenized into pieces, the model cannot "see" the individual letters to perform tasks like counting them.
A generic tokenizer trained on web text might break a medical term like "myocardial" into meaningless pieces. A specialized tokenizer trained on medical texts would keep it as a single unit. This preserves meaning, improves learning efficiency, and prevents wasting the model's limited context window on fragmented words.
Unlike older methods that use an "unknown" placeholder, modern tokenizers built on Byte-Pair Encoding (BPE) can represent any string. When faced with a new word, they don't crash or lose information. Instead, they break it down into the smallest known components, ensuring universal representability.
When an LLM sees the same document thousands of times, it prioritizes memorizing that specific content to lower its training loss. This gives a false impression of progress, while the model fails to generalize to new, unseen data. A study of Google's T5 corpus found one sentence repeated over 60,000 times.
If a document in the validation set has a near-duplicate in the training set, the model's high score is a lie. It's not demonstrating generalization; it's recalling something it has already seen. To prevent this, teams must deduplicate the entire dataset before splitting it into train and validation sets.
