Large, off-the-shelf multilingual models are often poor at less-common languages like Sanskrit due to data scarcity and improper tokenization. A smaller, focused team can outperform them by carefully curating a high-quality corpus, making domain-specific 'care' the key differentiator over raw computing power.
For languages like Sanskrit, a single visual character is often composed of multiple Unicode code points. Standard tokenization breaks these apart into 'orthographic shrapnel', forcing the model to relearn spelling. Splitting by grapheme clusters preserves the meaningful units, making invalid output unrepresentable.
The most significant challenge in building a Sanskrit GPT wasn't coding the transformer model but extracting clean text from PDFs. Issues like scanned images, legacy fonts, and encoding errors required more time than the AI development itself, showing that high-quality data sourcing is the primary obstacle and competitive moat.
