I would think a historian would love to have the data we have available now for any epoch.
so are historians preparing to use and collect and sort all that massive data available via internet and social media?
I can only talk about a project I was introduced to at the British Library during an Open Day of theirs which is their web archiving project. The British Library is attempting to preserve web resources of any '.co.uk' domain. An example they gave in their presentation was the archiving of campaign and political sites, which often disappear after the elections, but could be useful to historians and social/political scientists. They also are archiving tweets which can show trends during events with a big social media impact.
This is from the British Library: About the Web Archive
And they also have a blog which shows some of the projects using these archives already and also a bit more information about how it works: BL Blog
This is actually something that historians, and more enthusiastically and loudly librarians, archivists, etc., have been thinking about for some time now, and it is absolutely not clear that we have an answer. Although there are some attempts to catalog and preserve the web, it is not clear that these have the same long term stability that archives tend to have. Paper documents lasts a pretty long time as long as you keep them under good storage conditions and they never become inaccessible or unavailable due to changes in technology or software. Even when backwards compatibility is maintained many digital storage devices often can maintain data for only a span of 10-30 years and need to be maintained, whereas something like microfilm can last an order of magnitude longer. Both issues are a concern when it comes to whether or not these documents can or will be maintained over the long term.
By contrast to surviving paper records, digital evidence is much more liable to be taken down. Some organizations are cognizant of this and make regular backups of their archives precisely for the purpose of keeping a record (H-Net's discussion archives come to mind, but they are far from the only ones).
Roy Rosenzweig wrote an article about this some 12 years ago already. You can find it here He essentially argues that the problems we face are two fold. First is the actual preservation. Second is the, as your question suggests, actually using that data. Historians are very much accustomed to piecing together history via an incomplete historical record. No one is capable of personally sorting through the entire twitter archive, for example. Even a team of scholars working for years would have simply no chance at it. You might have a REASONABLE chance of trying to do something with that data only surrounding a specific event, but even then the amount of data exceeds many archive based research projects. The so-called computational turn of the Digital Humanities (although I would argue that DH far exceeds just the computational turn - this is a point of contention among those who do DH) offers one way of approaching it - using software to analyze the data. But that still leaves the question of interpreting that data, and here we are still vary much in the infancy of this kind of work. Some DH scholars argue that the computational turn means the death or at least deemphasis of theory, but I think most of our work still lies ahead of us as historians when it comes to a theory that will allow us to draw useful and insightful historical conclusions based on the analysis of large sets of data that emerge from things like social media.
I recently finished up writing a Master's thesis on commemorative social media in Iraq and Syria, so I might have a few things to contribute from a Middle East Studies angle. I was using all sorts of social media, and trolling various public and not-so-public corners of the internet for martyrdom and commemorative videos from the conflict.
There's a few takeaways from my experience. One is that much of the really interesting material for historians will be audio/visual, as well as text (which in a lot of ways mirrors classic written 'archive' material, but in a lot of other important ways is quite different). In my experience it's very difficult to figure out what bits of video are relevant to whatever you're studying or wanting to figure out. Tagging is generally useless, and data timestamps are up to the whims of whatever device is being used, meaning that actually placing a recording definitively in a certain place or time is very difficult and will only get more difficult as time goes on. At the moment, at least in my field, there's some easy workarounds. For example, working with this huge archive of video from the U.S occupation of Iraq, I got quite good and telling the year it was made based on the quality of film. Recordings from early in the conflict were generally sourced from old-fashioned tape recorders, which of course now have been replaced by cell-phone videos (I wrote a whole chapter about these). So there will be ways to date videos in this way in the future, but I suspect that we've reached a peak in video quality in that it will be quite difficult for non-technical historians of the future to place or contextualize digital visual material from one decade to the next.
In this sense, there's very little that's different from more traditional modes of archival research, that is, trolling through documents or microfilm or whatnot and trying to figure out its context in the situation that it's not easily placeable.
There are quite a few digital archivalists such as archive.org which keeps a record of all the jihadist material that gets floated around. There's a particularly good reason for this which speaks to another problem with digital material. Despite all the chat of the internet being forever, it's really not. Much of the material I'm interested in, that is Jihadi or militant material from certain periods, is actually quite difficult to find since so many groups have a vested interest in getting rid of it. Right in the middle of my research into jihadi material, Anonymous decided to shut down a huge amount of ISIS' media and supporter network. Laudable, but made my life quite difficult. Moreover, much of the material that will be particularly interesting to historians, such as journals or personal records, aren't kept in a centralized database, and might be lost forever.
Initiatives like the British Library's web archiving project is really interesting but points to another problem. As much as us English speakers like to think that our language is the language of the internet, other languages are increasingly being used, sometimes in quite novel ways (I'm thinking about Arabic here). Working with non-latin script online is still a nightmare, as anyone who has tried to use Arabic in word will know. Perhaps I'm getting off topic here.
Good question!
To add onto this, are there any current efforts to preserve or mine data from recent world events such as the Arab Spring?
Like many archives, a great deal of social media and internet data are being preserved in idiosyncratic ways, depending on the initiative of archivists, librarians, and scholars. Washington University, for example, is attempting to preserve social media responses to the unrest in Ferguson a year ago through the Documenting Ferguson project. You can read an interview with one of the people involved with the project here: http://www.processhistory.org/?p=843
On a larger level, I know that the Library of Congress has been working with Twitter to preserve its digital archive. They've been talking about that since 2010. I imagine the incoming Librarian of Congress will have a big role in deciding when and how that's made available. A recent update on that is available here: http://www.politico.com/story/2015/07/library-of-congress-twitter-archive-119698.html
Follow up question: are there any emerging fields that use software, machine learning, other tools to tackle these issues? I do both software and history (not historian at this point), would be interesting to see how I could apply myself to work in both fields.
As a medievalist, I find myself salivating over the sheer amount of material that modern/future historians have on this period to play with, then then I remember how much of said material is questionable at best.
I know the Library of Congress and the British Library are actively saving websites, I think it's safe to assume that other countries are doing the same.
I wonder about this all the time because I am a history enthusiast who works in digital advertising and see the sheer amount of data and meta-data we're dealing with.
The challenge is not even going to be recording everything. The challenge is going to be drawing analysis and research from such huge data sets which basically requires quant. people.
So my follow up question is: are there moves in the field to start incorporating more quantitative disciplines/people?
Doesn't social media merit the same treatment as "private correspondence" - i.e. not really worth preserving in its entirety?
I think it's a bit too recent a challenge for a historical perspective, but that's what I think.
I have heard (and please do feel to correct me) that actually instead of a abundency of data in the future we may actually struggle to find useful information due to the rapidly changing ways we access our information.
Take for example floppy discs. It has been just 15 years (at a guess) since they have been relevant and anything on them is harder to access than print material now. Having relevant hardware and software that will allow us to access the information contained within is just going to get harder. It is difficult to predict or future proof digital information. We may get a point where HTML or other software's become obsolete and we loose vast swaths of information.
Please do feel free to critique this, it is only based on a conversation I've had with archivists.
Side question: Given that microfilm is one of the more efficient ways to physically store semi-readable information in analog format, does any service out there print digital PDFs directly to microfilm?