Software / tool for archival research in a team?

by airportakal

I am currently doing a project in which I study the history of a certain building and it's related association in my town. I have collected a number of relevant archival documents, and skimmed through them for highlights of information.

This weekend, I will be going through them more thoroughly, but will be doing this with help from others. Currently, the scans of the documents reside in a number of Google Drive folders. I make annotations and partly transcribe the texts in a separate Drive document. Then I aggregate the most relevant info from these annotations and collect them in another document, which is not organised by archival document, but by subject (e.g. persons, construction work, etc).

I was wondering if there is a more efficient way of organising this research. Specifically with regards to the annotations, it is quite annoying to have a couple of scans separately from a word processing document, as it is difficult to link specific words, sentences and findings to specific parts of the scan.

Can you recommend tools for archival annotation, and are there any that work online so they can be used for collaborative research? Any others tips would be very welcome as well. It's the first time I do archival research. :-) Thank you!

silverappleyard

I don’t know how much has changed in the last three years, but there’s some discussion of tools in this old thread with u/restricteddata and u/Achaemenid-Empire that might be helpful.

whatafinebeerthisis

A search for "software" on r/AskHistorians brings me to this thread 4 months late!

Currently, the scans of the documents reside in a number of Google Drive folders. I make annotations and partly transcribe the texts in a separate Drive document. Then I aggregate the most relevant info from these annotations and collect them in another document, which is not organised by archival document, but by subject (e.g. persons, construction work, etc).

Brace yourself for this: Google Drive could have done a LOT of that transcription for you automatically! Image files and non-OCR'ed PDFs are automatically OCR'ed (including cursive handwriting). You just have to tell Google Drive to do that somewhere in preferences.

OP, are you or were you using a Mac for this project? If so, is/was the majority of your material saved as PDF or as images? I ask because if you were using a Mac and taking a bunch of smartphone images or uploading page scans, there's a great tool in Quick Look (the logo of an eye -- usually in Finder). Open quick-look and within it is a "Markup" icon. Once in markup, you can highlight text, underline, box, add arrows, stars etc. Just don't do this on top of text, because I presume that will screw up the ability for the text to be OCRed and therefore searchable.

Your use of Google Drive is extremely smart for two reasons: 1. you can set the drive to automatically OCR all content that isn't OCR'ed already, including image files (it even OCRs cursive handwriting alarmingly well!), and 2. it's a fabulous platform for sharing and collaboration. I rely extremely heavily on the ability to right-click any file in my google drive folder, from within Finder (not online), and select "share with google drive" to link words to sources in various documents. I also rely heavily on Apple Finder's tagging system. Unfortunately, I've reached a point where this setup is no longer viable however.

In the near future, I'll be migrating every single file I've every collected over to DEVONthink Pro. DTP falls in the category of "referencing software." I think it's been around for something like 25 years now. Software in this category is typically nicknamed a "bucket app." Here's a thread where people compare their favorite among these. DTP does OCR and has built-in AI ("machine learning") assets. I don't own a license but am playing with the free version right now. It's incredible.

You can set-up bots via an add-on called DEVONagent (also costs money) which automatically logs into library accounts, does fancy searches, and stores what it finds in a special folder -- crazy! Another add-on helps draw relations graphically. These would be hidden connections attained from information in the DTP database. This software's best asset is that it uses a lot of AI ("machine learning") to help speed-up the annoying "tagging" process for each bit of info you gather. Unfortunately, it's not cheap software -- but they have a very attractive student discount.

ronniethelizard

I work in software. At work I have a tool called confluence which allows me to create and edit documents and then share them between people. I can also attach files (e.g. photos, powerpoints, pdfs) to documents.

The company's website is: https://www.atlassian.com/

They have 2 other tools that may be of use:

  • JIRA: which allows you to create tasks, assign them to people, track progress, ...
  • Bitbucket: this allows you to keep a massive revision history of a document and share it between a large number of people. In addition, each person can create branches on documents make edits and then merge them together at the end. That said, this is specifically targeted at text files (though you can store other file types in a bitbucket repo).

I think you can rent space on their cloud to store things (so effectively you could communicate around the world fairly quickly).

Note: These tools are designed for software development so IDK if they will work well for you.

If you are looking for something free, github might work for you. Bitbucket mentioned above wraps around a software package called git which github uses. Like Bitbucket it is targeted at text files (though you can store other file types in git).

​

Another tool: latex. This allows you to create PDFs in a programmed manner so you can create a file with annotations in it and link to image files and then compile everything into a PDF document.

That said, this is notoriously tricky to use well (and I have heard it described as "programming a report"). If you have ever written HTML, this is a much more complex equivalent.