TanglishCode-SwitchingTamil AI
Tanglish: Tamil-English Code-Switching and AI
Tanglish is how many Tamil speakers really talk. Learn how Tamil-English code-switching works, why AI struggles with it and why models must handle it.
Eelam Lab · · 7 min
Tanglish is the everyday mix of Tamil and English that many Tamil speakers use, from a single English word dropped into a Tamil sentence to Tamil typed entirely in Latin letters. It follows real patterns rather than being "broken" Tamil, and because it is so common in speech, texting and social media, any AI system that wants to understand Tamil speakers has to handle it well. This guide explains how Tanglish works, why it confuses AI and what a Tamil model needs to get it right.
What is Tanglish?
Tanglish (sometimes spelled Tamglish) is a blend of Tamil and English. Linguists describe this kind of mixing as code-switching: moving between two languages within a conversation, a sentence or even a single word.
You will hear Tanglish in Chennai offices, Toronto family kitchens, Singapore hawker centres and WhatsApp groups everywhere. A typical sentence might be:
Naan office-ku late-aa vandhen. ("I came to the office late.")
The grammar is Tamil. The word order is Tamil. The suffixes -ku ("to") and -aa (an adverb-forming ending) are Tamil. But "office" and "late" are English. That is Tanglish in a nutshell.
Wikipedia has a short overview of Tanglish as a phenomenon, though usage varies a great deal from person to person and place to place.
Tanglish is not "bad Tamil"
It is easy to dismiss Tanglish as laziness or a loss of language skill. Research on code-switching in many languages suggests otherwise: bilingual speakers switch in systematic, rule-governed ways, and switching often signals identity, humour, closeness or topic.
For many Tamil speakers, especially in cities and the diaspora, Tanglish is simply how they talk. A grandparent in Jaffna and a teenager in Scarborough might both use English words in Tamil sentences, just different ones in different ways. Treating that speech as an error to be corrected means ignoring a large share of real Tamil.
That does not mean formal Tamil is less important. Many people value pure Tamil and use it with care. A good AI system should understand both, and respond in the register the user prefers.
The main forms of Tanglish
Tanglish is not one thing. It shows up in several distinct forms, and each creates different challenges for AI.
1. English words inside Tamil grammar
This is the most common pattern. English nouns, adjectives or verbs slot into a Tamil sentence and take Tamil suffixes:
- "Meeting-la irukken." ("I'm in a meeting.")
- "Weekend-la movie paakalaam?" ("Shall we watch a movie at the weekend?")
English verbs are often paired with the Tamil helper verb பண்ணு (paṇṇu, "do"):
- "Cancel pannunga." ("Please cancel it.")
- "Naan call panren." ("I'll call.")
This is a very productive pattern. Almost any English verb can join the language this way.
2. Switching between whole sentences
Some speakers move between full sentences in each language:
"I told him already. Avan kekkala." ("I told him already. He didn't listen.")
3. Romanised Tamil
Many people type Tamil using Latin letters because it is faster on a phone keyboard. This is still Tamil, just written in a different script:
"Saapteengala?" ("Have you eaten?")
There is no single standard spelling. One word might be written several ways. வணக்கம் (vaṇakkam, "hello") may appear as "vanakkam", "vanakam" or "vannakkam". The letter ழ (ḻa) is written as "zh", "l" or "z" depending on the person, which is why you see both "Tamizh" and "Tamil".
4. English words in Tamil script
The reverse also happens. English words get written in Tamil script, such as மீட்டிங் (mīṭṭiṅ, "meeting") or பஸ் (pas, "bus"). Text might also mix scripts in one line:
நான் meeting-ல இருக்கேன். ("I'm in a meeting.")
Why Tanglish confuses AI
Most AI systems are built on an assumption that each sentence belongs to one language. Tanglish breaks that assumption in several ways.
Language identification fails
Many pipelines start by detecting the language of a text and routing it accordingly. A sentence like "Naan call panren" may be labelled English (because it uses Latin letters) or flagged as unknown. The model then responds in the wrong language, or not at all.
Spelling is inconsistent
Because romanised Tamil has no fixed spelling, the same word appears in many forms. A model needs to see enough real examples to learn that "poren", "porean" and "pōṟēṉ" may all point to போறேன் (pōṟēṉ, "I'm going").
Tokenizers split mixed text badly
Tokenizers trained mostly on English or mostly on formal Tamil may split Tanglish into awkward fragments. A word like "office-ku" mixes an English root with a Tamil suffix, and a tokenizer needs to have seen such forms to handle them efficiently. We cover this in how a Tamil tokenizer works.
Speech recognition stumbles
When people speak Tanglish, English words are pronounced with Tamil sounds and rhythm, and they take Tamil endings. Speech recognition trained on either pure English or formal Tamil reading often mishears them or forces them into the wrong script.
Training data is mostly "clean"
Large datasets are often filtered to keep text that looks like one language. That filtering can quietly remove mixed text, so models see less Tanglish than people actually use. Tanglish also overlaps heavily with spoken Tamil, which is already underrepresented compared with written Tamil. See spoken vs written Tamil for more on that gap.
Why a Tamil model must handle Tanglish
If an AI system ignores Tanglish, it effectively ignores how a great many Tamil speakers communicate day to day. That has practical and cultural costs.
- Usefulness. People type the way they talk. A Tamil assistant that cannot read "naalaikku meeting irukka?" ("is there a meeting tomorrow?") is not very helpful.
- Inclusion. Diaspora speakers, who may be more comfortable in Tanglish than formal Tamil, should not be shut out of Tamil technology. Our article on the Tamil diaspora explores this further.
- Keeping Tamil in use. If Tamil tools only accept formal Tamil, many people will simply switch to English. Meeting speakers where they are can keep Tamil in daily use.
- Respecting choice. Understanding Tanglish does not mean replying in it by default. A good model should understand mixed input and still be able to answer in formal Tamil, spoken Tamil or English, as the user prefers.
What it takes to handle Tanglish well
At Eelam Lab, we are building Nila (நிலா, nilā, "moon"), a Tamil large language model trained from scratch. Handling Tanglish well means paying attention at every stage:
- Collect real mixed speech and text. Natural conversations, not just scripted sentences, so the model hears how switching actually happens.
- Keep mixed text in the data rather than filtering it out as "noisy".
- Train the tokenizer on both scripts, including romanised Tamil and English words in Tamil script.
- Use consistent transcription guidelines for recordings, so English words in Tamil speech are captured faithfully.
- Evaluate with real speakers who use Tanglish, from different regions and generations.
Nila is still in development, and how well it understands Tanglish will depend on how much natural, mixed speech it learns from.
Frequently asked questions
Is Tanglish a separate language?
Not in the usual sense. Tanglish is a way of mixing Tamil and English, usually with Tamil grammar as the base. It varies a lot between speakers, regions and generations rather than having one fixed form.
Should I avoid Tanglish when contributing my voice?
No. If you naturally mix English into your Tamil, speak the way you normally do when talking freely. That natural speech is exactly what AI models are missing. When you are asked to read a specific Tamil sentence aloud, just read it as written.
Can AI translate Tanglish into formal Tamil?
It can, in principle, if it has been trained on enough examples of both. Today most tools struggle, because they have seen little Tanglish and few pairs of mixed and formal text.
Is romanised Tamil the same as Tanglish?
Not exactly. Romanised Tamil is Tamil written in Latin letters, and it can contain no English at all. Tanglish refers to mixing the two languages. In practice the two overlap a lot, especially in texting.
Talk the way you really talk
The best way to teach AI about Tanglish is to let it hear real people. In the contribution studio, you can talk freely for up to five minutes per clip about anything: your day, a family story, how you cook a favourite dish. Mix in English if that is how you speak. You can also type or paste text you have written yourself. Every contribution needs explicit consent and is reviewed by people before training. Record your voice.
Add your voice to Nila
Record a sentence, tell a story in your own dialect, or share something you have written. It takes a few minutes.
Open the contribution studio →