I was recently watching a documentary about the Hittites and they talked about how Hattusha contained thousands of tablets with a completely unknown language, eventually researchers were able to decode the language.
This reignited a question I have had for a long time; how do researchers even begin to understand completely lost, ancient languages? I would assume they try and look for patterns or structures from already understood languages, but a more detailed explanation would be appreciated. Thank you
A cryptography book I have explains how they managed to read the cuneiform script of the Achaemenid Dynasty. Researchers found a lot of cuneiform writing in the ruins of Persepolis, but unfortunately no one can understand it, since the writing has been out of fashion since 100 BCE.
However, at the start of the 19th century, a German schoolteacher by the name of Georg Friedrich Grotefend noticed that there are repeated phrases in the script; he guessed that it corresponds to "X, king of kings" or "Y, son of King Z". He knows the names of the Achaemenid kings of the era thanks to Herodotus' works; therefore, he simply substituted the cuneiform letters with the phonetic sounds.
To make the example clearer, here's a screenshot from my book.
From there, they try to work out what the other letters are, and eventually the cuneiform script was "cracked" the the point where we can now read the Epic of Gilgamesh written in its original language.
This is also how Egyptian hieroglyphs were cracked; the Rosetta stone consists of three parts. From top to bottom, it's the same decree written in hieroglyphs, demotic and Greek.
Speaking specifically of Hittite, we had already managed to decipher cuneiform scripts--Hittite's version of cuneiform was derived from Old Assyrian cuneiform with a smaller number of symbols. Of course, deciphering the script itself doesn't get you all the way, since Hittite did some strange-looking things. It didn't consistently show the voicing distinction Assyrian did (between [p b t d k g]), and doubled vowels and consonants in various places (probably to show length or accent).
The language was successfully deciphered and identified once we had a substantial corpus of texts and a Czech linguist named Bedřich Hrozný undertook a pretty extensive analysis of them. Because the nature of the texts was largely formulaic (e.g. diplomatic correspondence and inventories), it was possible to identify recurring structures, and Hrozný managed to identify regular correspondences between Hittite and other Indo-European languages. Importantly, he also identified a Hittite consonant with a proposal of Swiss linguist Ferdinand de Saussure. De Saussure, about 2 decades earlier, had proposed 3 coefficients sonantiques (today called laryngeals, under the assumption that they were h-like), sounds that had either merged with other sounds in the history of Indo-European or simply been lost. These sounds were proposed to handle tricky correspondences like English father, Latin pater, Sanskrit pitar, which were until then unexplained. Importantly, these proposed sounds had no direct attestation in any Indo-European known at the time of de Saussure's writing. Hrozný identified in Hittite cuneiform the consonant ḫ, which appeared exactly where de Saussure had proposed we would find 2 of these laryngeals.