Hey all,
Long time lurker, I'm a computer programmer looking for some guidance and insight. Etymology in particular has always fascinated me. It's interesting being able to trace a word back to its roots, or see how latin and greek combined and then evolved.
There are a number of software tools that can measure how different two pieces of text are and return a number, e.g. soundex & string distance. This is one library of tools that can be used to analyse text and otherwise manipulate it.
I'm on the hunt for objective analysis using tools like these and others which mathematically compare measurements in the differences in a reference text that has evolved over time.
Are there any good sources of specific texts that have changed throughout history but have a clear and distinct transformation alongside a common language? I would like to try my hand at using these tools to analyse phonetic / stem / string distances over time as well. I'm thinking the Bible might be one such example?
I understand I may come across as vastly underestimating the complexity of language evolution as there are no doubt breaks in continuity, sudden surges of change followed by periods of relative stability too, not to mention completely new words come and go, so I am also keenly interested in learning about all of this.
My goal is to be able to work in reverse - take a modern text and degenerate it toward a faux-ancestral language (or multiple ancestors). Input x years and use key metrics to guide the de-evolution. The end result doesn't need to be 100% accurate to the real-world evolution, so I wouldn't expect to put in english and get fluent ancient roman, but I would like to be passably close for a casual observer beyond simply randomising phonemes.
Thanks in advance, and if there's a more fitting form for these questions happy to be pointed in the right direction.
Point of clarifiation: Are you looking for infinite translation software? Or are you wanting to type in (say) modern English text and get something like Chaucer Doth Tweet?
Either way, in wanting to look at different editions of a text over time, you're kind of reinventing the linguistic wheel. Even a glance over the "history of X language" Wiki page will show you that scholars have mapped out the general evolution of diphthongs, consonant combinations and silent letters, extent of declension use, etc already. (This is enormously complicated by local dialects and lack of standard spelling in the Middle Ages/early modern era). Your best bet is going to be familiarity with the language in question.
P.S. "Ancient Roman" is called Latin.
I'm afraid you're looking at reinventing something rather more complex than the wheel all at once.
First, there's already a technique for reconstructing a language's or language family's ancestors: it's called the comparative method. While there are now programs that will do the phylogenetic aspect of the comparative method, grouping languages according to a list of (human-selected and human-coded!) characters, there's not yet anything out there that does the actual work of reconstructing proto-forms. That's all people, and has been since the late 18th century.
There's also a key misunderstanding here (at least as I understand your goal of creating "olde-worlde" languages), that older languages are somehow different in kind from modern languages. The fact of the matter is that languages are languages are languages, and the idea that old languages "look" a certain way while modern languages don't is a matter of sampling bias.
If you really want to understand how the sound structures of languages change, pick up a (couple of) introductions to phonology and historical linguistics. Campbell's and Crowley and Bowern's intros to historical linguistics (both in their fourth edition) are excellent textbooks. Of course, that's just the start, but you're not going to do it a priori by coding some program.