Dev.to AI 🤖 Ai 👁 0 📖 5 min read

How well do AI models read Tanglish? I tested 13 of them

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked I'm a Tamil speaker. When Tamil people text, we mostly type Tamil words in English letters and throw in English words wherever it

How well do AI models read Tanglish? I tested 13 of them

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I'm a Tamil speaker. When Tamil people text, we mostly type Tamil words in English letters and throw in English words wherever it is easier. That mix is called Tanglish. "Naalaikku kaalaila ezhara manikku bus" is a normal message. In Tamil script the same thing is நாளைக்கு காலைல ஏழரை மணிக்கு பஸ், and very few people bother to type that on a phone.

AI models get tested on Tamil in proper Tamil script. So I wanted to check one thing: if I give a model the same sentence in Tanglish, how much worse does it get? I'm calling that drop the script penalty.

The benchmark has 60 everyday comprehension questions. Each one exists in three versions with the same meaning: English, colloquial Tamil in Tamil script, and Tanglish. The Tamil script and Tanglish versions use the same words, so the script is the only thing that changes between them. The four answer choices are in English in every version.

Here is one question in all three forms:

English: The bus is at seven thirty tomorrow morning. Mum said I have to be at the stand half an hour before that. What time is the bus?

Tamil script: நாளைக்கு காலைல ஏழரை மணிக்கு பஸ். அதுக்கு அரை மணி நேரம் முன்னாடியே ஸ்டாண்ட்ல இருக்கணும்னு அம்மா சொன்னாங்க. பஸ் எத்தனை மணிக்கு?

Tanglish: Naalaikku kaalaila ezhara manikku bus. Adhukku ara mani neram munnadiye stand la irukkanum nu amma sonnanga. Bus ethana manikku?

A) 6:30 AM B) 7:00 AM C) 7:30 AM D) 8:30 AM

The 60 questions fall into six groups:

  • 17 on idioms and slang whose literal meaning misleads. "Alwa kuduthutan" has nothing to do with sweets.
  • 15 on number, time and kinship words.
  • 9 on grammar forms, such as the hearsay ending "-aam".
  • 8 on words that become ambiguous in Roman letters. "Arai" can be a slap, a half or a room.
  • 7 short message threads.
  • 4 on tone and sarcasm.

Each model is allowed to think aloud and has to end with a line that says "Answer: X". Code checks that final letter. There is no judge model.

Two early mistakes shaped the design. My pilot asked for a bare letter, and Claude Haiku lost points for showing its working even when it reached the right answer. I also had a question that needed subtraction, and two models got it wrong in English, which meant it was testing arithmetic. I changed the scoring and took the sums out.

Models Tested

I ran 13 models on Kaggle's free quota:

  • Google: Gemini 3.7 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite, Gemma 4 31B and Gemma 4 26B
  • Anthropic: Claude Sonnet 5 and Claude Haiku 4.5
  • OpenAI: GPT-5.4 mini, GPT-5.4 nano and the open gpt-oss-20b
  • GLM-5 from Z.ai and Grok 4.20 (non-reasoning) from xAI

I leaned toward small and cheap models on purpose. An app that serves Tamil users at volume will run a Flash-Lite or a nano, so that is where a script penalty would reach real users. The larger models are there as a reference line.

Some models are missing. Kaggle lists Grok 4.5 and Grok 4.6 but returned "model not found" for both. gpt-oss-120b, DeepSeek-R1 and Qwen 3 Next timed out or were overloaded during my runs. GPT-5.5 hit a quota reservation error on one of the three tasks, so it has no complete score yet.

The final run was 2,456 model calls and used $1.97 of quota.

Findings

Model English Tamil script Tanglish Change
Gemini 3.7 Flash 100% 100% 100% 0
Gemini 3.5 Flash 100% 100% 100% 0
Claude Sonnet 5 100% 98% 100% +2
Gemma 4 26B 100% 98% 100% +2
Gemini 3.1 Flash-Lite 100% 100% 98% -2
Gemma 4 31B 100% 100% 98% -2
Gemini 3.5 Flash-Lite 100% 98% 98% 0
GLM-5 100% 100% 95% -5
GPT-5.4 mini 100% 98% 93% -5
Grok 4.20 (non-reasoning) 100% 98% 93% -5
Claude Haiku 4.5 100% 93% 80% -13
GPT-5.4 nano 98% 77% 77% 0
gpt-oss-20b 100% 88% 70% -18

Accuracy per model in Tamil script and in Tanglish

English is the control. Every model scored 98% to 100% on it, so a miss in the other two columns comes from the Tamil.

Google's models barely noticed the script. Gemini 3.5 Flash and Gemini 3.7 Flash got all 180 prompts right. The two Flash-Lite models and both open Gemma 4 models missed at most one question in any version, and Claude Sonnet 5 was level with them. I expected the small Gemma models to struggle and they did not.

Five models did pay a penalty. GLM-5, GPT-5.4 mini and Grok 4.20 each lost 5 points going from Tamil script to Tanglish. Claude Haiku 4.5 lost 13 and gpt-oss-20b lost 18, on sentences they had mostly understood in Tamil script.

GPT-5.4 nano scored 77% in both scripts, so its trouble is with colloquial Tamil itself and the script makes no difference.

Accuracy by question type in Tamil script and in Tanglish

Number, time and kinship words did most of the damage. Across the 11 models that missed anything, that group fell from 92% in Tamil script to 79% in Tanglish. The bus question above is the clearest case. Six of the 13 models got it wrong in Tanglish, reading "ezhara" (half past seven) as 6:30 or 7:00, and only one got it wrong in Tamil script. "Mundhaa naal" (the day before yesterday) tripped five models in Tanglish, and four of them answered "yesterday".

Idioms were easier than I expected, at 95% in both scripts. Every model knew that "alwa kudukradhu" means stringing someone along.

Two results went the other way, and I can't explain them. "நாளன்னைக்கு" (the day after tomorrow) was read as "tomorrow" by six models in Tamil script, Claude Sonnet 5 among them, and by three in Tanglish. "பயங்கரமா", said about a biryani the speaker ate three plates of, was taken literally as "frightening" by four models in Tamil script and by one in Tanglish.

Those two aside, the misses lean one way. Sixteen questions were missed only in Tanglish, and one was missed only in Tamil script.

There are limits to how far I'd trust the small gaps. Each model ran once, so a difference of one or two questions is noise. Sixty questions is a small set, and 25 of them were answered correctly by every model in every version. The Tamil is the colloquial kind I know, and other regions speak and spell differently.

Next I would test heavier SMS spelling with dropped vowels, add regional dialects, and repeat every run so the small gaps can be trusted. I would also like scores for the models Kaggle could not serve this week.

My Benchmark

Benchmark: Tanglish Script Penalty on Kaggle

The three tasks behind it:

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.