Why do Tamil and Korean share so many cognates and other shared words

by thenewwwguy2

I find this interesting and I was discussing it on a language learning subreddit because we were discussing language origins and how they evolved. I speak tamil and was once speaking to my grandfather in it at the Seoul Airport when once of the security guards said that I speak very good Korean. Initially I thought this was just funny and a coincidence, but I was surprised to see that there are numerous words which are similar or identical between the two languages, including the words for I, Mom, Dad, Grass, Fight, etc.

I can’t rule out that this is coincidence, and upon some reading (the wikipedia article for the Dravidian-Korean language theory, lol), it seems like people have noted and proposed similarities but no historical link exists to justify the theory’s accuracy.

So are these similarities simply coincidental, or is there some historical link between Dravidian India and Korea that exists? Moreover, if the link does exist, why and how did these languages develop similarly? If it doesn’t, why do they have this much coincidence?

iwaka

The similarities are coincidental. It is not unusual to see chance similarity between unrelated languages, especially in shorter words.

The way we compare languages in historical linguistics is by using the comparative method. The comparative method is quite simple to explain, but it takes a lot of time and training to learn to use it properly, because it's very easy to abuse its power and get weird results, like a link between Korean and Tamil. The method itself works by looking at regular and recurrent sound correspondences. Let's break that down.

Regular means that in two or more languages, the correspondence is always the same. So for example, English /θ/ corresponds to German /d/ in word-initial position:

English German
three drei
thunder Donner
thistle Distel

This is regular, and we would expect this correspondence to hold for all cognates between the two languages. You can't have /θ/ corresponding to /d/ in one word, but then /f/ in another without an adequate explanation.^1

Recurrent means that the correspondence should occur multiple times within the dataset. This is usually the biggest problem with long-range comparisons like the one you mentioned. All too often, many "correspondences" appear only once, in which case they might as well be completely random.

Another very important detail is that all sounds have to be accounted for. Quite commonly in these fringe theories, they only compare a part of a word and just disregard the rest. That is a huge methodological flaw, and greatly increases your chances of finding a connection where there isn't one (which I guess is the point for people who abuse the comparative method like that). If you say there is a fused affix in the root, you must prove that the affix indeed existed and was productive. You cannot just lop off half a word (or even a single sound) willy-nilly, that is not how the comparative method is used.

Another thing pertaining to the Korean-Tamil hypothesis but also quite common in other such comparisons. Tamil is uncontroversially part of the Dravidian language family. When comparing language families at a higher level, we only compare data from reconstructed proto-languages, and not from daughter languages. Why? Because there are 80 languages in the Dravidian family (according to Glottolog), and that hugely increases the probability of seeing a chance resemblance if you can just cherrypick words from any language in the family. Imagine if you're comparing two families with over 1000 languages each, you'll have chance resemblances left and right. So if a language is demonstrably and incontrovertibly part of a larger linguistic family, you compare your data to the proto-language, not the daughter languages.

Lastly, semantics. You have to be really careful where word meanings are concerned, because that's another way to hugely increase your chances of finding random noise and attributing meaning to it. When you're first comparing two languages (or language families), you look for words with identical meanings, which also have regular and recurrent correspondences in all of their sounds. Only after a plausible connection is established and the sound correspondences worked out, you can allow for a limited amount of semantic shift, assuming that all sounds correspondences are regular. If you skip this step, you're allowing almost anything to be compared. You should see some of the mental hoops people jump through to justify a connection. Not too long ago when answering a similar question on r/linguistics, I went through a linked article where the author was comparing the words "salmon" and "to fly". Their explanation? Well see, when salmon spawn, they travel upriver, and sometimes jump through shallows, and this jumping gives the appearance of flying. So basically, if you do not control for semantics, it's a free-for-all where anything goes.

Real historical linguistics is boring. It relies on meticulous analysis and most new discoveries are pretty "duh!". There are people who cannot resist the temptation of being the first to make a great discovery, and do not mind abusing the comparative method to do it, which of course invalidates all their efforts.


  1. You can have different correspondences in different environments (e.g. word-initial, between vowels, after a stressed vowel, etc). Sometimes you come across seeming exceptions, which may be explained through lexical borrowing.