ThirukkuralTamil LiteratureLanguage Data

Thirukkural and AI: What an Ancient Text Teaches

What the Thirukkural, Tamil's celebrated book of couplets, teaches us about language data, careful transmission, accuracy and building AI for Tamil.

Eelam Lab · · 7 min

The Thirukkural is a classical Tamil work of 1,330 short couplets on ethics, statecraft and love, traditionally attributed to the poet Thiruvalluvar. Scholarly dates for it range from around 300 BCE to about 500 CE. For people building AI, it is a useful lesson in how language survives: through careful transmission, trusted commentary, and communities who care about getting every word right. It also shows, quite sharply, what today's AI systems still get wrong about Tamil.

This is not an article about whether an ancient poet "predicted" artificial intelligence. It is about what a text that has been read, recited and argued over for centuries can teach a lab working on Tamil language data today.

What is the Thirukkural?

திருக்குறள் (Tirukkuṟaḷ, often written Thirukkural or Tirukkural) is one of the most widely read works in Tamil. Its name combines திரு (tiru, "sacred" or "revered") and குறள் (kuṟaḷ, "short", the name of its two-line verse form).

Its structure is remarkably regular:

  • 1,330 couplets (kurals), arranged in 133 chapters of exactly 10 couplets each.
  • Three books: அறம் (aṟam, "virtue"), பொருள் (poruḷ, "wealth" or "polity") and இன்பம் (iṉpam, "love", also called காமம், kāmam). The three books have 38, 70 and 25 chapters respectively.
  • One metre throughout: the kural veṇpā, with four metrical feet in the first line and three in the second.

The date of the Thirukkural is genuinely uncertain. Scholarly estimates vary widely, from a few centuries BCE to around 500 CE, and the question is still debated. Very little is known for certain about Thiruvalluvar himself.

The work has been translated into many languages, and English translations go back to the 19th century, including G. U. Pope's well-known version of 1886. The Wikipedia article on the Kural is a reasonable place to start reading more.

Lesson 1: Density is not the same as volume

Each kural is just seven metrical feet. In that space, Valluvar compresses whole arguments. The opening couplet reads:

அகர முதல எழுத்தெல்லாம் ஆதி பகவன் முதற்றே உலகு

(akara mutala eḻuttellām āti / pakavaṉ mutaṟṟē ulaku)

A common rendering: "As the letter A is the first of all letters, so the Primal Being is first in the world."

Modern AI culture tends to equate quality with quantity: more data, more parameters. The Thirukkural is a reminder that a small body of text can carry enormous meaning. The full work is only around 1,330 couplets, far too small to train a language model on its own, yet it shapes how millions of people think and speak.

For a Tamil model, this matters in two ways. First, Tamil does not need to match English word-for-word in data volume to be represented well; it needs data that is clean, varied and genuinely Tamil. Second, a model has to handle compressed, allusive language without flattening it into generic paraphrase.

Lesson 2: Transmission is a chain of care

The Thirukkural reached us because people kept copying it, reciting it and teaching it. For centuries, Tamil texts survived on palm-leaf manuscripts, which decay and had to be recopied by hand. Every copy was a chance for an error to creep in, and every careful scribe was a chance to catch one.

Data for AI works the same way. Text is scraped, cleaned, converted, deduplicated and fed into training. At each step, Tamil can be damaged: characters broken by bad encoding, vowel signs detached from consonants, or text silently replaced with machine translation. The Unicode Tamil block defines how Tamil script is stored digitally, and a surprising amount of web Tamil does not follow it cleanly.

We explore the technical side of this in our article on the Tamil tokenizer. The principle is the one the manuscript copyists understood: if you do not care for the text at every step, the errors accumulate.

That is why every contribution to Eelam Lab is reviewed by people before it is used in training. It is slower, and it is worth it.

Lesson 3: Commentary is part of the text

Almost no one reads the Thirukkural without help. Medieval commentators, the most famous being Parimelaḻakar in the 13th century, wrote explanations that shaped how the couplets were understood for generations. Modern readers rely on school notes, scholarly editions and the explanations of teachers and parents.

For language data, this is a lesson about context. A couplet alone is not enough; the understanding around it, often passed on orally, is part of what makes it meaningful. When an elder explains a kural in their own spoken Tamil, that explanation is valuable data in its own right. It connects classical language to everyday speech, a gap we discuss in spoken vs written Tamil.

Lesson 4: Accuracy is a form of respect

Ask many general-purpose chatbots to quote a specific kural, and you may get a confident answer that is subtly wrong: a misremembered word, a mismatched number, or an invented couplet that sounds plausible. For a text that Tamil speakers know by heart, this is not a small bug. It is a sign that the system does not really know the language.

The Thirukkural itself has advice on this. Kural 423, in the chapter on wisdom (அறிவுடைமை), says:

எப்பொருள் யார்யார்வாய்க் கேட்பினும் அப்பொருள் மெய்ப்பொருள் காண்ப தறிவு

(epporuḷ yāryārvāyk kēṭpiṉum apporuḷ / meypporuḷ kāṇpa taṟivu)

Roughly: "Whatever you hear, from whomever you hear it, wisdom is seeing the truth of the matter."

A Tamil model should be built in that spirit: trained on verified text, honest about uncertainty, and careful with sources that communities hold dear. That is a goal for Nila (நிலா, nilā, "moon"), the Tamil large language model Eelam Lab is building from scratch. Nila is still in development, and we would rather it say "I am not sure" than misquote Valluvar.

Lesson 5: Classical text is not the whole language

It would be easy for a Tamil AI project to lean heavily on classical literature. Public-domain works like the Thirukkural are well digitised and beautifully written. But a model trained mostly on classical and formal text would speak like a textbook, and would barely understand a conversation in a Jaffna kitchen or a Coimbatore bus.

The Thirukkural lived for centuries alongside ordinary speech, proverbs, songs and stories. A good Tamil dataset needs that same balance: classical and modern, written and spoken, Tamil Nadu and Sri Lanka and the diaspora. You can read more about that variety in our guide to Tamil dialects.

A kural for learners and builders

Kural 391, the opening couplet of the chapter on learning (கல்வி), is often quoted to students:

கற்க கசடறக் கற்பவை கற்றபின் நிற்க அதற்குத் தக

(kaṟka kacaṭaṟak kaṟpavai kaṟṟapiṉ / niṟka ataṟkut taka)

"Learn thoroughly, without flaw, what is worth learning; then live in keeping with it."

It is hard to think of a better summary of responsible data work. Choose what is worth learning from. Clean it carefully. Then build something that behaves accordingly.

Frequently asked questions

Who wrote the Thirukkural?

It is traditionally attributed to Thiruvalluvar (திருவள்ளுவர், Tiruvaḷḷuvar). Almost nothing is known about his life with certainty, and much of what is told about him is legend. Scholars continue to debate both his identity and the date of the work.

How many couplets are in the Thirukkural?

There are 1,330 couplets, organised into 133 chapters of 10 couplets each, across three books on virtue, wealth and love.

Is the Thirukkural used to train AI models?

As a public-domain classical text, it can appear in training data for Tamil models, but it is far too short to train a model on its own. Its real value is as high-quality, carefully transmitted text that a good Tamil model should be able to quote and explain accurately.

Why do AI chatbots misquote the Thirukkural?

Most large models see relatively little clean Tamil text, and much of it is noisy or machine-translated. Without enough reliable examples, they produce plausible-sounding but incorrect Tamil. Better, carefully reviewed Tamil data is a large part of the fix.

Help carry Tamil forward

The Thirukkural survived because people chose to keep it. Everyday Tamil, your family's stories, your grandmother's proverbs, the way your town says "yes", needs the same care.

You can contribute in the studio by recording your voice (reading a sentence or talking freely for up to five minutes), typing text, or uploading a document you own or that is in the public domain, such as an out-of-copyright book or your own handwritten notes. Every contribution needs your explicit consent and is reviewed by people before training.

If an elder in your family can explain a favourite kural in their own words, that recording is a small piece of living Tamil. Record it in the contribution studio.

Add your voice to Nila

Record a sentence, tell a story in your own dialect, or share something you have written. It takes a few minutes.

Open the contribution studio →