[Methodology] How will historians deal with primary sources on computers?

by Kikool42

So my question is more about the methodology of History than actual fact. I hope it's not going to be deleted because it aims just a little bit about the present and the future. I'm really asking about the methodology and the way historians work. Note that I will talk using future tense but what I'm talking about might be already happening.

So, I'm used to seeing records on papers of archives, letters written by thinkers and authors, and so on. I have myself a book of Nietzsche with his first letters and his first written texts. Historians deal with a lot of primary sources written on paper.

But now that we have come to an electronic era... are you still going to work the same way? How will you collect the equivalent of letters written between people (e-mails, Facebook conversations perhaps) ? Isn't there any risk of false authenticity (could be easily modified), or any risk that the amount of records we have decreases with time because documents and sources are easier to lose on computers because they can be deleted?

I guess what I'm asking is how will historians deal with primary sources that were written directly on computers. Letters, texts, even art, all that made on computer today. How do you deal (or will you) with documents or sources made by the geniuses of today considering they are on computer and might not be as "fancy" as paper material (considering the slang words of the Internet and the fact that communications on Facebook or Skype will not be as fancy as 19th century hand written letters) ?

Do you think the fact that we're moving into a more virtual and technological era will change the methodology and the approach of historians on today's sources?

Thank you in advance, and sorry if the question is a bit off timeline.

caffarelli

First off, historiography and “method of history” questions are encouraged here!

How will you collect the equivalent of letters written between people (e-mails, Facebook conversations perhaps)?

I personally will collect them along with my colleagues, because I am an archivist! This saving and authenticity work is the realm of the archivist, not the historian, and in the digital age it still will be the work of archivists, not historians. Archivists save and tend to the records, historians use them to make the history. So the rest of the stuff like how on earth are they going to sift through all this stuff, not my problem, but how to save it, yes my problem.

The public static web is pretty easy to save so long as you have servers to burn, just look at archive.org. Emails are kinda tricky to save, but not impossible. The leading archivist in the field of email preservation is Chris Prom. Here’s a quick overview of his methods of saving email. Email I’m not personally worried about, email will make it through history just fine if people want it to.

Facebook unfortunately has just about no way to harvest/scrape it, outside of some very limited methods of exporting your own content. Facebook owns all that content and they have locked it down tight as a drum against traditional web archiving methods. I’ve done a few experiments on trying to save a few student political organization Facebook pages (I worked on this maybe 4 months ago) and about all I could get out of it was some MHTML files that were mostly poop. Facebook is high-risk for loss, unless they do something like Twitter’s groundbreaking agreement to save data with the Library of Congress back in 2013.

Isn't there any risk of false authenticity (could be easily modified), or any risk that the amount of records we have decreases with time because documents and sources are easier to lose on computers because they can be deleted?

I'm not sure how digitally literate you are so I'll try to keep this ELI5 level, let me know if you want more gory details. Basically archivists are on top of this too. Once we get a file in our pale spidery nerdfingers it’s not going to be modified. There’s some problems with the fact that we get most records when someone dies or they have passed out of use for an organization (right now most files I work with are from the 90s-00s) so they’re 10-20 years old when we get them, but we have a lot of ways to work with corrupted files. After that, we use something called checksums to monitor file integrity, you can google them, they’re not very complicated. It basically adds up all the data in a file and makes a number; six months later you add it up again and if it’s the same number, it’s the exact same file. Different number, it’s degraded or corrupted. The Signal from LoC has good content on this. If you look into products like DSpace they do this babysitting for us. Industry standard is double or triple backup on any archival server.

So yeah, I can keep the records authentic and safe, but historians are on their own for the rest of the work. :)

mpsellus

Hi - this is a fascinating issue.

Agree with you on the problems related to Internet correspondence. However, for webpages (including blogs, social media posts), there is already some form of archival system that future historians will find useful, called the Wayback Machine, which makes a copy of every webpage every few months (or in even shorter intervals, upon request).

Here is the actual site... https://archive.org/

... and an interesting New Yorker article on it. http://www.newyorker.com/magazine/2015/01/26/cobweb

82364

Adding on, do you anticipate a lack of self-curation (i.e., information being spread across tens of thousands of emails) being a problem?