Tamil AITokenizationUnicode

Tamil Tokenizer Explained: How AI Reads Tamil Text

How does a Tamil tokenizer turn text into tokens? Grapheme clusters, Unicode, BPE and why English-first tokenizers waste Tamil, explained with examples.

Eelam Lab · · 8 min

A Tamil tokenizer is the component that splits Tamil text into small units, called tokens, before an AI model can read it. A good Tamil tokenizer keeps the letters a reader actually sees (grapheme clusters) intact and learns common Tamil word parts, while a tokenizer built mainly for English often shreds Tamil into many tiny byte-level fragments. That difference affects cost, speed, how much text a model can handle at once and how well it learns Tamil grammar.

What is a tokenizer, and why does it matter?

Language models do not see letters or words directly. They see a sequence of numbers, and each number stands for a token. A token might be a whole word, part of a word, a single character or even a fragment of a character's underlying bytes.

The tokenizer decides how text gets cut up. That decision shapes everything that follows:

  • Cost and speed. Models do work per token. More tokens for the same sentence means more compute.
  • Context length. Models can only look at a fixed number of tokens at once. If Tamil takes three times as many tokens as English, the model can "remember" a third as much Tamil.
  • Learning. If tokens line up with meaningful pieces of the language, such as roots and suffixes, the model has an easier time learning grammar.

For English, popular tokenizers are well tuned. For Tamil, they often are not.

How Tamil is stored on a computer

To understand Tamil tokenization, start with how Tamil text is encoded.

The Tamil script is an abugida

Tamil script combines consonants and vowels into single visible letters. Traditionally the alphabet is described as 12 vowels (உயிர், uyir, "life"), 18 consonants (மெய், mey, "body"), 216 combined vowel-consonant letters (உயிர்மெய், uyirmey) and one special letter, ஃ (āytam), for 247 in total.

The combined letters are not stored as single characters. Instead, Unicode stores a base consonant followed by a vowel sign. The Tamil Unicode block runs from U+0B80 to U+0BFF, and you can see the full Unicode Tamil code chart on unicode.org.

One visible letter, several code points

Take கு (ku). It looks like one letter, but it is two code points:

  • க (U+0B95, ka)
  • ு (U+0BC1, vowel sign u)

Or take the word வணக்கம் (vaṇakkam, "hello"). It is seven code points:

வ + ண + க + ் + க + ம + ்

The ் here is the புள்ளி (puḷḷi, "dot"), a mark that removes the inherent vowel from a consonant. A reader sees five units: வ, ண, க், க, ம். Those visible units are called grapheme clusters, and the rules for finding them are set out in Unicode Standard Annex #29. Under those rules the Tamil puḷḷi attaches to the consonant before it rather than joining it to the next one, so வணக்கம் is five grapheme clusters. (Unicode 15.1 added a conjunct rule that keeps clusters such as Devanagari क्ष (kṣa) together, but up to and including Unicode 18 it does not cover Tamil, whose puḷḷi is not classed as a conjunct linker.)

Then come the bytes

Most text is saved as UTF-8. Every Tamil code point takes three bytes in UTF-8, while basic English letters take one. So:

  • "hello" is 5 bytes.
  • வணக்கம் is 7 code points, which is 21 bytes.

Keep that number in mind. It matters a lot for what comes next.

How most AI tokenizers work: BPE

The most common approach is byte pair encoding (BPE). In simple terms:

  1. Start with the smallest units, often individual bytes.
  2. Look through a large training corpus and find the pair of units that appears together most often.
  3. Merge that pair into a new token.
  4. Repeat until you reach a target vocabulary size.

Frequent sequences end up as single tokens. In English, common words like "the" or "language" usually become one token each because the tokenizer saw them millions of times.

A related method, the unigram language model used in tools like SentencePiece, works from the other direction: it starts with many candidate pieces and prunes them. Either way, the outcome depends on the training corpus.

Why English-first tokenizers waste Tamil

If a tokenizer was trained on data that is mostly English, it has spent nearly all of its merges on English patterns. Very few merges are left for Tamil.

Byte fragments instead of letters

Byte-level tokenizers can always represent any text, which is a strength. But if they have learned few Tamil merges, they fall back to small byte sequences. In the worst case, வணக்கம் could become up to 21 separate byte tokens, while "hello" is typically one token.

Worse, merges can cut across a grapheme cluster. A token might contain the end of one letter and part of the next, or split a consonant from its vowel sign. The model then has to learn, from scratch, that these fragments belong together.

High "fertility"

Researchers often measure tokenizer fertility: the average number of tokens per word. A higher number means the language is being broken into more pieces. Tamil, being agglutinative, already has long words. Combine long words with poor merges and Tamil can end up needing several times as many tokens as an equivalent English sentence. Exact figures vary a lot by tokenizer and text, so treat any single number with caution.

The practical result: Tamil users pay more (where pricing is per token), get slower responses, fit less text into the context window and receive lower-quality output. For a wider look at these problems, read why AI struggles with Tamil.

What makes a good Tamil tokenizer?

There is no single perfect design, but several principles help.

Train on lots of real Tamil

The most important factor is simple: the tokenizer must be trained on a large, varied body of Tamil text. That includes formal writing, spoken-style transcripts, dialects and mixed Tamil and English, so the learned pieces reflect how people actually write.

Respect grapheme clusters

A Tamil tokenizer should avoid splitting a visible letter in the middle. One way to do this is to pre-segment text into grapheme clusters and only allow merges between whole clusters. That way, கு is never split into க and a stray vowel sign.

Learn meaningful word parts

Tamil words are built from roots and suffixes. Consider:

  • வீடு (vīṭu, "house")
  • வீட்டில் (vīṭṭil, "in the house")
  • வீட்டுக்கு (vīṭṭukku, "to the house")

A well-trained tokenizer will often learn pieces like the root and common case endings such as -இல் (-il, "in") and -க்கு (-kku, "to"). Then a word it has never seen whole can still be represented with a few meaningful tokens. Notice, though, that the root changes shape (வீடு becomes வீட்டு), which is one reason purely statistical methods do not always line up with grammar.

Normalise text first

Some Tamil vowel signs can be stored in more than one way. For instance, the sign ொ (o, U+0BCA) has a canonical decomposition into ெ (U+0BC6) followed by ா (U+0BBE). Both look the same on screen. Without Unicode normalisation (typically NFC), a tokenizer might treat identical words as different strings. Older Tamil text in legacy, pre-Unicode font encodings also needs converting before it is usable.

Handle Tanglish and romanised Tamil

Many people type Tamil in Latin letters or mix English words into Tamil sentences. A Tamil tokenizer that only knows Tamil script will struggle with "naan office-ku poren". Including mixed and romanised text in tokenizer training helps. See Tanglish and code-switching for why this matters.

Balance vocabulary size

A larger vocabulary means fewer tokens per sentence but a bigger, more memory-hungry model. A smaller vocabulary means longer sequences. Choosing the size is a trade-off between efficiency and model capacity, and it should be tested on real Tamil rather than guessed.

How this connects to Nila

Eelam Lab is building Nila (நிலா, nilā, "moon"), a Tamil large language model trained from scratch rather than adapted from an English model. Training from scratch means the tokenizer can be designed for Tamil from day one instead of inherited. You can read more about the full process in what it takes to build a Tamil large language model.

A tokenizer can only learn the Tamil it is shown. If its training text is all formal news, it will be efficient for news and clumsy for a conversation in Jaffna Tamil or a family chat in Tanglish. That is why varied, real-world contributions matter even at this early stage.

Frequently asked questions

Why does Tamil use more tokens than English?

Each Tamil code point takes three bytes in UTF-8, Tamil words are often long because of suffixes, and most popular tokenizers were trained mainly on English. Together, these mean Tamil text is split into many more pieces than English text with the same meaning.

What is a grapheme cluster in Tamil?

It is a unit that a reader perceives as one letter, even if it is stored as several Unicode code points. For example, கு (ku) is stored as க plus the vowel sign ு but is read as one letter.

Does a better tokenizer make the AI smarter?

Not by itself, but it helps a lot. A good Tamil tokenizer lets the model see more Tamil per context window, spend less compute per sentence and learn word structure more easily. The model still needs good training data.

Should a Tamil tokenizer include English?

Yes, to some degree. Many Tamil speakers mix English words into Tamil or write Tamil in Latin letters, so a useful tokenizer needs to handle both scripts sensibly.

Help shape how AI reads Tamil

Tokenizers learn from real text, and the more varied that text is, the better they serve every kind of Tamil speaker. You can contribute in the contribution studio by typing or pasting Tamil text you own, uploading a public-domain document (PDF, Word, text or photos of pages, up to 25 MB), or recording your voice. Every contribution requires explicit consent and is reviewed by people before training. Contribute Tamil text.

Add your voice to Nila

Record a sentence, tell a story in your own dialect, or share something you have written. It takes a few minutes.

Open the contribution studio →