Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Unlike older methods that use an "unknown" placeholder, modern tokenizers built on Byte-Pair Encoding (BPE) can represent any string. When faced with a new word, they don't crash or lose information. Instead, they break it down into the smallest known components, ensuring universal representability.

Related Insights

A useful mental model for an LLM is a giant matrix where each row is a possible prompt and columns represent next-token probabilities. This matrix is impossibly large but also extremely sparse, as most token combinations are gibberish. The LLM's job is to efficiently compress and approximate this matrix.

A generic tokenizer trained on web text might break a medical term like "myocardial" into meaningless pieces. A specialized tokenizer trained on medical texts would keep it as a single unit. This preserves meaning, improves learning efficiency, and prevents wasting the model's limited context window on fragmented words.

A unified tokenizer, while efficient, may not be optimal for both understanding and generation tasks. The ideal data representation for one task might differ from the other, potentially creating a performance bottleneck that specialized models would avoid.

A common misconception is that Transformers are sequential models like RNNs. Fundamentally, they are permutation-equivariant and operate on sets of tokens. Sequence information is artificially injected via positional embeddings, making the architecture inherently flexible for non-linear data like 3D scenes or graphs.

Beyond the obvious lack of non-English training data, Large Language Models are architecturally biased. Their tokenization process, designed for English, inefficiently breaks down other languages into more fragments. This increases operational costs and reduces comprehension, creating a structural disadvantage.

Models like ChatGPT struggle with basic string manipulation because their fundamental unit of understanding is a "token," which can be a whole word or a sub-word. If a word like "Strawberry" is tokenized into pieces, the model cannot "see" the individual letters to perform tasks like counting them.

Autoencoding models (e.g., BERT) are "readers" that fill in blanks, while autoregressive models (e.g., GPT) are "writers." For non-generative tasks like classification, a tiny autoencoding model can match the performance of a massive autoregressive one, offering huge efficiency gains.

The 2017 introduction of "transformers" revolutionized AI. Instead of being trained on the specific meaning of each word, models began learning the contextual relationships between words. This allowed AI to predict the next word in a sequence without needing a formal dictionary, leading to more generalist capabilities.

Custom tokenizers and embeddings, created for a foundation model, can be repurposed to enhance other data engineering tasks. They can improve OCR accuracy on domain-specific documents, allowing for better text-based processing and avoiding the higher cost of vision models.

Due to how tokenization works, non-Latin based languages like Hindi, Thai, or Greek can require two to five times more tokens to represent the same amount of text as English. This creates a 'language tax,' making AI-powered applications significantly more expensive to operate for non-English users.