[META] Small analysis most popular questions AskHistorians

by Isinator

Some days ago I noticed Reddit has an API enabling people to extract Reddit data. For some time I've been interested in this subreddit and I decided to analyse some AskHistorians data. The result can be found here. It's nothing too in-depth, but I'm sure the data has more potential once you attack it from some interesting angles.

Edit: thanks for all the feedback, appreciated a lot. I'm definitely planning on reworking the analysis based on the comments provided (there's a lot of legitimate criticism). I'm very interested in what type of questions would be interesting to you, don't hesitate to let me know :).

Since this isn't really a question I added the [META] tag but I'm not too sure if this is a moderator thing only. Please remove this if I wasn't allowed to use it.

sunagainstgold

Thanks for this; it's terrific and so are you!

Georgy_K_Zhukov seems to be in another league than everyone else. Having made nearly a thousand comments in roughly 1/4 of all top questions asked by users is quite a feat. In no way I want to underestimate the work done by other users, it's just that there really is a gap of about 500 comments with the second contender.

Honestly, /u/Georgy_K_Zhukov deserves all the credit he can get and more for the work he puts into AskHistorians. It's great to see even just one part of that quantified so neatly.

some people seem to never sleep (sunagainstgold)

You're not wrong.

Georgy_K_Zhukov

So on the one hand, "HEY! LOOK AT ME!!!!" On the other though, I know I shouldn't be looking a gift horse in the mouth, but is it possible to rerun your analysis with some way to exclude distinguished 'Mod' comments? I feel that my #1 positioning is due primarily to my moderation comments. Not to say that I'm not writing answers as well, of course, but I would venture that the ratio is skewed to more mod comments than 'regular' comments, especially given the general prominence of mods in the top 20. I don't know what data was included in the 'pull' that you did, but if an indicator for Distinguished is one of them, I'd really love to see it re-run with them excluded, or else noted as such.

restricteddata

I suspect my posting frequency graph is distorted by the AMAs I have done — those big beacons that stick out.

My strong aversion to posting on Wednesdays is kind of amusing, especially when overlapped with my "time of day" posts. On Wednesdays I typically teach during the times of day I would otherwise be tempted to check on here.

AdamMonkey

Nice work. It confirms my believe that Roman history is very popular on this sub.

brigantus

It's a bit of a pity there's some overlap of username labels but I don't think there's an easy way to solve this issue and having the names on the graph itself is kind of nice.

There's a package that makes it pretty straightforward, ggrepel.

thedeliriousdonut

Woah. Huh, that's weird. I just met /u/yodatsracist recently and we were talking about reddit's algorithm and now here they are in a post about reddit's algorithm. I mean, not entirely about the algorithm, but yeah. Guess you start seeing people everywhere once you know them.

historianLA

I would actually switch the axes on the time since creation and length of answer graph. That would visualize the issue better since I think length is the dependent variable in this instance. That shows that the most thoughtful answers are not the first nor the late arrivals. They are relatively early but take time to produce.

bobobo1618

In case you weren't aware, you could collect most of this data from BigQuery without scraping the API. There's a dataset here: https://bigquery.cloud.google.com/table/fh-bigquery:reddit_comments.2015_05 has every comment/post from Reddit's inception until around August 2016 and makes it super easy to query them.

smayple

I wonder about another reason shorter comments have higher scores. Long comments, at least from what I see as a lurker, tend to include a lot of obscure bits of info that are beyond what a lay person like me tends to be able to put into context. Shorter comments tend to have less depth and address the question at hand in a more focused manner, which is easier to understand. I think most of this subs subscribers are probably not professional historians

Erpp8

When you mapped answer length vs. Score, did you include only answers, or all comments? Because that could explain the negative correlation. A lot of top comments are either mods reiterating rules, or interesting follow-up questions, both of which are quite short and quickly accumulate points.

Halinn

People become like ancient war time know Roman years?

jofwu

 I can see two reasons how this could be the case...

My guess would be that people aren't normally patient enough to read (and then vote on) long answers.

nemtrif

Thanks for the analysis!

Personally, (and I am just an amateur historian) I've had a bit of a problem with the "comprehensive, in-depth" rule for the sub-reddit. In practice, it seems that the moderators favor walls of text even if they don't even answer the question asked. Like it or not, some questions are best answered by short responses and these are discouraged by the culture on this subreddit.

grapp

the word cloud affirms some of my own instincts about which of my posts will likely get traction and which won't

Tiako

Interesting that /u/yodatsracist, /u/vertexoflife and I are the only really old timers on the top twenty. I wonder if removing mod comments would change that. Also you can see what month I got a new job, talk about a life in one chart.

jschooltiger

This is really cool stuff. Not to pile on at all because several people have already mentioned this, but it would be interesting to get the info without the distinguished comments. As a flair with a somewhat obscure field, I'm sure that a lot of mine that are counted are mod comments for post removals, rules reminders, etc. So I would love to see it with that teased out.

Thanks for doing this, it's really cool.

gobberpooper

Maybe this is for another thread, but do you think we're getting a very strong bias in answers because a lot of the visible answers end up coming from the same 10 or 15 people? So that rather than getting answers from a wide range of the historical community at large, we're getting answers primarily through the lens of commiespaceinvader, sunagainstgold, yodatsracist etc? Not that I'm contesting their merit in any way, or the work they've done, I'm just curious if anyone else sees this.

heygivethatback

Meta-comment for a meta-post: how exactly did you extract the data? Would you be open to posting some useful links for people who are familiar with R (looks like you used R for your graphics?) but unfamiliar with API's?