How do historians/archaeologists/linguists go about reconstructing dead languages?

by saddetective87

What if they have no audible reference?

keyilan

There are a number of factors that contribute to our overall understanding of now-lost languages. Written records help, when they exist, and sometimes we're even lucky enough to have figures like Pāṇini or traditions like that of Tang China where people had enough interest to actually write down their own analyses of their language and its earlier forms.

However the most significant tool we have to work out reconstructions is what's called the comparative method. I'm quoting here from Campbell (1999), which is cited at the end of this comment. Campbell defines the comparative method as follows:

The aim of reconstruction by the comparative method is to recover as much as possible of the ancestor language (the proto-language) from a comparison of the descendant languages, and to determine what changes have taken place in the various languages that developed from the proto-language.

To do this, we take the currently spoken ancestors of the parent languages and looking for correspondences. It's not limited to that though. For example now-lost pronunciations in pre-modern Korea for words borrowed from Chinese are also an important factor, and contribute to the data which we analyse.

Continuing from Campbell (1999):

By applying the comparative method to related languages, we can postulate what [a known] earlier [common] ancestor was like … [C]omparing English with its relatives, Dutch, Frisian, German, Danish, Swedish, Icelandic and so on, we attempt to understand what the protolanguage, in this case called 'Proto-Germanic', was like. Thus, English is, in effect, a much-changed 'dialect' of Proto-Germanic, having undergone successive linguistic changes to make it what it is today, a different language from Swedish and German and its other sisters, which underwent different changes of their own.

He adds one point that I'll take issue with:

Therefore, every proto-language was once a real language, regardless of whether we are successful at reconstructing it or not.

The issue here is that generally the Proto-Language are the reconstruction itself, and at least in my neck of the woods (Sinotibetoburmanwhatever historical lingusitics), we do not assume that the reconstruction actually represents an actually spoken language. Let's say you reconstruct A and B as two proto-sounds. A later became X, and B later became Y. Maybe A turned into X before B ever developed, so that A and B never actually co-existed in the same speech variety at the same time. However in the reconstruction you'd possibly account for both A and B, because — again in my subfield — reconstructions are correspondences. We'll come back to that word, so keep it in mind. The reconstructed proto-language doesn't therefore represent an actual singular dialect at a point int he past, but rather a set of correspondences between all living and attested descendants. I'm not 100% and someone like /u/rusoved who knows his Indo-European can correct me on this, but it's my understanding that in some circles the reconstruction is actually considered a real once-living language.

Either way, that particular difference in how it's viewed isn't important. Both tell us the same thing: Here's how today's language developed in relation to its sister languages.

So now back to correspondences. There are a few incredibly important requirements that need to be satisfied for a reconstruction to pass muster. Correspondences working out is a big part of that. By "working out" what I mean is that, if you have English "ship", Dutch "schip", German "Schiff", Danish "skib" — all words that mean the same basic thing and in languages that are known to be related — It's not enough to say "oh yeah ok the P in English is like F in German". Yes, that's true, but you need to be able to account for it beyond just "hey this looks like that over there". You have to be able to show that the corresponding sounds are productive-predictive, which just means productive and predictive but it sounds cooler with the hyphen doen't it?

One really famous example of this is Grimm's Law (one of the Brothers Grimm), which is super interesting but which I'll leave you to Wikipedia for the details. Moving on.

So how do correspondences past this test of being productive-predictive? The following is an example using a dialect of Japanese (from Tōkyō) and a dialect of Ryūkyūan (Shuri). This is taken from Vovin (2011), once again cited at the end. I'm using these examples because I already typed them up in a Reddit-friendly format a couple days ago. So here are the words:

meaning Tōkyō Japanese Shuri Ryūkyūan
this kore kuri
bird tori tui
rare mare mari
rice kome kumi
place tokoro tukuru
temporary kari kai

Let's assume we don't know that these are related languages, but we think they might be. We need to find the correspondences. In this case we can see that some of the correspondences are e#→i#, ri#→i# and o→u. The # symbol is marking boundaries, and to the left of the arrow is Tōkyō Japanese.

So, to show that these langauges are likely related, we need to make sure the correspondences are productive-predictive. That is, having worked out what the correspondences are, you should be able to look at kore and be able to deduce that the corresponding Shuri word is kuri even if you didn't have that word in front of you. If I pull an imaginary example of tokome as a Tōkyō word, you should be able to come up with tukumi as the Shuri equivalent (those are probably made up words; I didn't bother to check if they mean anything, it's just an example). If then we find out that Tōkyō tokome does in fact match up with a Shuri word tukumi, then we know that our correspondences are sound (heh).

The whole point of this is that it's not enough to say "hey obrigado in Portuguese looks like arigato in Japanese". If you think those are likely related, you need to work out the correspondences. In the case of "obrigado", it's 100% a coincidence that those words kinda look similar. You won't be able to find productive correspondences.

The bigger issue — and the reason you can spend your entire life on this kind of task of working out reconstructions — is that if you really want your reconstruction to be good, it has to not only predict a newly discovered dialect that might appear one day, having avoided attention until now (that's the predictive part), but it also has to explain how the changes happened historically in a way that's justified based on what we know of how language changes.

For that, you need to know something about how other languages in the world have changed, historically. The first test is solid correspondences. The second test is being able to justify those correspondences as plausible.

A huge part of the hard-science side of linguistics is in laboratory phonology, acoustical phonetics, and basically doing a lot of fine measurements on speech data to be able to explain with anatomical certainty how the sounds are happening the way they are, and how the articulatory production and cognitive reception of those sounds influences how they are likely to change. We know why Southern Mandarin and Cantonese speakers are starting to not distinguish between /n/ and /l/ (see here for more on this particular change). We know because we can explain it physiologically as a consequence of the anatomical machinery used to make those two sounds.

And we can see in languages like Norwegian and Swedish changes that have been happening in Shanghainese, for development of retroflex consonants on the former and vowel backing/raising with the latter. Shanghainese and Swedish are not at all related, but the same sorts of things that cause sound change in one are happening in another. That's a good thing if you're looking at a third language and find the same kind of changes, because there's well-documented cases showing the same sort of thing.

In other words, for the historical linguist, all of these minute analyses are being done with reference to every attested language that the researcher can get their hands on to see if a specific shift has happened elsewhere in the world (which it usually has).

I'm sorry that got super long. I've summarised some of my other answers to similar questions here, but if you want anything more in depth, let me know.

References:

  • Campbell, Lyle (1999) Historical Linguistics: An Introduction. MIT Press, Cambridge.

  • Vovin, Alexander (2011) Why Japonic Is Not Demonstrably Related to Altaic or Korean. Historical Linguistics in the Asia-Pacific region and the position of Japanese ICHL, vol. 20.