Tamil AILarge Language ModelsNila

Tamil Large Language Model: Building One From Scratch

What does it take to build a Tamil large language model from scratch? Data, tokenizers, dialects, consent and evaluation, explained in plain language.

Eelam Lab · · 8 min

A Tamil large language model is an AI system trained to read, write and understand Tamil, ideally as a first language rather than as an afterthought. Building one from scratch means gathering a large, varied and consented body of Tamil text and speech, designing a tokenizer that respects the Tamil script, and testing the model against how Tamil people actually speak and write. This article walks through each of those steps and explains why Eelam Lab chose the from-scratch path for Nila (நிலா, nilā, "moon").

What is a Tamil large language model?

A large language model (LLM) is a neural network trained to predict the next piece of text given everything that came before it. Do that over enough text and the model starts to pick up grammar, facts, style and even some reasoning. You can read a general overview on Wikipedia's large language model page.

Most well-known LLMs are trained mainly on English, with other languages making up a small share of the data. They can produce Tamil, but often it reads like a translation: stiff, overly formal, sometimes grammatically off, and unaware of regional speech.

A Tamil large language model flips that balance. Tamil is the centre of the training data, the tokenizer is designed around Tamil script, and the evaluation asks "does this sound right to a Tamil speaker?" rather than "is this roughly correct?".

Fine-tuning versus training from scratch

There are two broad ways to get a model that handles Tamil.

Fine-tuning an existing model

You take a model already trained on mostly English data and continue training it on Tamil. This is cheaper and faster, and it can work reasonably well for some tasks. The downside is that the model's foundations, including its vocabulary and much of its internal knowledge, were shaped by another language. Tamil gets squeezed into a structure that was not built for it.

Training from scratch

You start with an empty model and train it on data where Tamil comes first. This is harder and needs more careful data work, but the model learns Tamil patterns directly instead of through an English lens. Nila, the model Eelam Lab is building, is being trained from scratch for exactly this reason. It is not a fine-tune of an English model.

Neither approach is "wrong". But if the goal is a model that treats Tamil as a native language, with its own rhythm, idioms and dialects, starting from scratch gives the most control.

Step one: data, and lots of variety

Every LLM is a reflection of its training data. For Tamil, the challenge is not just quantity but variety.

Tamil is spoken by tens of millions of people across India, Sri Lanka, Singapore, Malaysia and a global diaspora (estimates put the total at more than 80 million when second-language speakers are included). It has one of the longest continuous literary traditions of any living language, with the earliest Sangam poetry usually dated to somewhere around 300 BCE to 300 CE, though scholars debate the exact range.

A good Tamil training set needs to reflect all of that:

  • Written Tamil: books, essays, news, letters, school material and public-domain archives.
  • Spoken Tamil: everyday conversation, which differs a lot from the written form (it uses different verb endings, pronunciation and vocabulary).
  • Dialects: Jaffna, Batticaloa, Chennai, Madurai, Kongu, Tirunelveli, Malaysian, Singaporean and diaspora speech all have their own vocabulary and sounds.
  • Culture: recipes, festivals, proverbs, kolam patterns, folk songs and oral history, the kind of knowledge that rarely appears on the open web.
  • Mixed speech: Tamil and English blended together, often called Tanglish, which is how many people actually talk. See our guide to Tanglish and code-switching.

Much of the Tamil that exists online is news, film content and formal writing. Grandmothers telling stories in Batticaloa Tamil, or a family in Toronto switching between Tamil and English at dinner, are almost absent. That gap is why Eelam Lab collects contributions directly from volunteers through the contribution studio.

Scraping whatever text you can find is the fastest way to build a dataset. It is also how you end up with copyright problems, low-quality machine-translated pages and content people never agreed to share.

For Nila, every contribution requires explicit consent. Contributors must own what they submit, or it must be in the public domain. Contributors are also told up front that contributions may be licensed to third parties. Contributions are reviewed by people before they are used in training, which helps catch errors, duplicated content and material that should not be there. Privacy follows the principles of Canada's PIPEDA, and you can read the details on our privacy page.

Quality review matters more for Tamil than for English. A badly encoded page, a mistranscribed recording or a block of machine-translated text can teach the model the wrong patterns, and with less total data available, each mistake carries more weight.

Step three: a tokenizer built for Tamil

Before a model sees any text, a tokenizer chops it into small units called tokens. Tokenizers designed around English often split Tamil words into many tiny fragments, sometimes even breaking a single visible letter into separate pieces.

That has real costs. More tokens per sentence means the model uses more compute for the same meaning, fits less Tamil into its context window and has a harder time learning word structure. A Tamil large language model needs a tokenizer trained on Tamil text that keeps grapheme clusters (the letters a reader actually sees) together and learns common Tamil word parts.

We go into detail in how a Tamil tokenizer works.

Step four: handling Tamil grammar

Tamil is agglutinative: words are built by stacking suffixes onto a root. A single word can carry what English spreads across a whole phrase.

For example, வீட்டிலிருந்து (vīṭṭiliruntu, "from the house") combines வீடு (vīṭu, "house") with markers for "in" and "from". Verbs pile on tense, person, number and negation in the same way.

This means a Tamil model sees an enormous number of distinct word forms, many of them rare. A model has to learn the building blocks, not just memorise whole words. Good data and a good tokenizer do most of the heavy lifting here. For a deeper look at why this trips up general-purpose AI, read why AI struggles with Tamil.

Step five: speech as well as text

A lot of Tamil life happens out loud. Elders who never wrote much down hold vocabulary, proverbs and stories that exist nowhere in print. Dialects are often easier to capture in speech than in writing, because written Tamil tends to drift toward a standard form.

That is why Eelam Lab collects voice recordings as well as text. In the contribution studio, you can read a Tamil sentence aloud or talk freely for up to five minutes per clip. Recordings help with speech recognition and also preserve how Tamil really sounds across regions.

Step six: evaluation that Tamil speakers trust

Benchmarks for Tamil are limited compared with English, and many are translated from English, which carries English assumptions with them. A useful Tamil large language model needs evaluation that checks:

  • Grammar and natural phrasing, judged by fluent speakers.
  • Whether the model understands dialect words and spoken forms.
  • Cultural knowledge: festivals, food, literature, history.
  • Handling of mixed Tamil and English input.
  • Whether it avoids flattening every dialect into one "standard" voice.

Human judgement from Tamil speakers in different regions is essential. No automated score can fully tell you whether a sentence sounds like something your aunt in Jaffna or your friend in Madurai would say.

Why build Nila at all?

When a language is underrepresented in AI, its speakers get worse tools: weaker search, clumsier translation, voice assistants that do not understand them and chatbots that answer in awkward textbook Tamil. Over time, younger speakers, especially in the diaspora, may simply switch to English because the technology works better in English.

A Tamil large language model built by and for the Tamil community is one way to push back. It treats Tamil as a full language worth building for, not a translation target. Nila is still in development, and its quality will depend directly on the breadth of Tamil it learns from.

Frequently asked questions

Is Nila a fine-tuned version of an English model?

No. Nila is being trained from scratch, with Tamil at the centre of its training data. That lets its vocabulary and internal patterns be shaped by Tamil itself rather than adapted from another language.

Why can't existing AI tools just learn more Tamil?

They can improve, and many have. But when Tamil is a small slice of the training data and the tokenizer was designed for English, there is a ceiling on how natural the Tamil output becomes. A dedicated Tamil model can go further on dialects, spoken Tamil and culture.

What kind of data does a Tamil LLM need?

It needs written and spoken Tamil, formal and informal registers, many dialects, mixed Tamil and English speech, and cultural knowledge such as proverbs and recipes. Variety matters as much as volume, because a model only learns the Tamil it sees.

Can I remove my data after contributing?

You can request deletion by emailing privacy@eelamlab.com, and contact details are deleted within 30 days. Data that has already been used to train a model cannot be removed from that model, but it is excluded from future training.

Help build a Tamil model that sounds like you

If you speak Tamil, in any dialect and at any level of fluency, your voice and writing can help. Visit the contribution studio to record a short clip, tell a family story, share a recipe or upload public-domain text (PDF, Word, text or photos of pages, up to 25 MB). Every contribution is consented, reviewed by people and used to help Nila learn Tamil as it is really spoken. Ten minutes from you is ten minutes of Tamil that might otherwise never reach an AI model. Start contributing.

Add your voice to Nila

Record a sentence, tell a story in your own dialect, or share something you have written. It takes a few minutes.

Open the contribution studio →