Search programs for Historians

by SillyCrawlingThings

Hi historians of reddit,

​

(new account due to doxxing etc.) I have started my PhD just yesterday. The subject falls within the confines of the history of religion. I am confident in my abilities surrounding the religion part, but the history part not so much (good thing I have some years to work on it).

​

My concrete question to you is: what sort of programs do you use to search through documents (or perhaps even to find documents in online databases as well)? I will have to, among other things, trace the history of a few specific terms, so a program in which I can fill in whole lists of search terms and which gives me an analysis based on this would be great (specifically within pdf and other text files). I guess my google searches do not contain the right key words as I only find pdf-finders and things like windows search files.

The1Brad

Could you explain further? Do you want a list of digital archives? Are you looking for a tool to search exclusively digital archives? Or are you looking for a text mining/ data analysis tool?

If you're looking for digital archives for your particular field, I'm sure a lot of people around here could help you. Just make a new post with a clearer prompt. If you're looking for a digital archive search engine, there are a number out there and usually the archives themselves have a search engine, but I don't think there's a collective historical archive search engine, outside of Google Scholar, which is terrible. I would love if someone created something competent but I don't think it will happen any time soon.

If you're looking to see how many times a word appears in published documents for a given year or over the course of time, check out Google Analytics. I'm not familiar with its usage, but I've seen historians chart the usage of terms like "the United States are" versus "the United States is" over time to answer historical questions.

If you are looking for an analysis tool like Google Analytics but just for a particular set of documents, I think you have to input the datasets and create the program this yourself (I'm not an expert in this but I think you could just modify someone else's program). It's been awhile but I've seen this used to create interesting charts and graphs in the Benjamin Franklin Papers Project, The Texas Slavery Project, and The Stephen F. Austin Papers Project. The University of Virginia has a complete collection of Digital History Projects that may also be of use.

I hope this helps answer your question. Sorry if it doesn't.

restricteddata

I don't know of any off the shelf tool that does what you suggest. It's a non-trivial thing to do for PDFs, even with good OCR. This is likely the kind of thing you'll need to build yourself.

Fortunately, this is not as hard as one might think. The Programming Historian includes some very nice tutorials on how to do various forms of text-mining in Python, such as this one, this one, this one, this one, and so on. There's no one way to do it; there are several approaches and programs that might be of use. Poking around the lessons might give you some ideas. And note that while learning to program might seem daunting, it is a highly marketable skill (including in history programs), and not nearly as hard as it at first seems (this is all under the category of "scripting" and usually uses high-level languages with robust libraries, so you're not doing the nitty-gritty aspects of programming and can focus more on the logical flow).