Why is geometric mean used to estimate war death tolls?

by [deleted]

I was reading this wikipedia entry of list of wars by death toll and found it odd that the geometric mean is used to estimate the death toll over the arithmetic mean without any explanation. Now, I understand that the geometric mean is smaller than the arithmetic mean and it would make sense to have a more conservative number as the death tolls might often be enumerated to higher numbers for political reasons. But, it also makes the number more susceptible to outliers which are small in magnitude. Say, you had three records of 1000, 9000 and 11000. In this case, the AM of 7000 is likely to be closer to the actual number than the GM of 4626.

lcnielsen

If you read carefully, you will see

Note 1: The geometric mean is the middle of the quoted range, taken by multiplying together the endpoints and then taking the square root.

In your example, the 9000 would not have an impact on the geometric mean, which would instead be c. 3300. To understand why this is done, consider what you're trying to do - you're trying to give a single value indicative of a range, not the average of some number of "records". Even if you were to have some number of records, there's no inherent reason to prefer an estimate that has been recorded many times, unless you have multiple independent methods that yield the same result. Instead, what you're typically trying to do is aggregate estimates with upper and lower bounds. If you're making an estimate of one aspect, you're going to be multiplying together factors like base population, increased mortality, duration of conquest, etc. You're going to have some confidence range for each factor, let's say you're pretty sure the correct value for any one factor is between 1.5x and 1/1.2x of your point estimate, and very sure it's between 3x and 1/2.4x. Let's say you're multiplying together five factors. Then, at the extremes of every factor you will get the ranges:

X * (1/(1.2^5 ) - 1.5^5 ) = X*(0.4 - 7.59)

X * (1/(2.4^5 ) - 3^5 ) = X*(0.012 - 243)

Let's say X = 1000, then your low-confidence range will be 400 - 7590, and your high-confidence range 12 - 243000. In the real world you'll also have some bounds that'll allow you to impose further constraints, but that's unimportant at this moment.

Now, how do you express these incredibly unwiedly estimates with a single number? Obviously you can't take the arithmetic mean, that's not going to make any sense (you can't have the low-confidence estimate be 4000 and the high-confidence one be 120 000!). But if you take the geometric means of the ranges, you'll get about 1700 for both the high- and low-confidence range. As it turns out, this ends up being a pretty good number.

Why? This is informally known as a Fermi problem. You have five random variables multiplied together: X = x1 * x2 * x3 * x4 * x5. Note that log(X) = log(x1) + log(x2) + log(x3) + log(x4) + log(x5), and incorporating the error factor E for each variable, we have log(X) + log(E^5 ) = log(X) + 5 log(E). But by the mean value theorem central limit theorem (I must be tired today...), when we add together normally distributed errors (which we can reasonably take the logarithms to be, since our range of uncertainty is multiplicative*), the total error only increases as the square root of the number of errors; hence our expected range is only about log(X) +-2.2*log(E) or X * E^(2.2 ) .

*The power of the CLT is that the individual errors don't actually even have to be normally distributed, as long as they are independent, the total error will still be normally distributed - but that's besides the point here!

This will reduce our narrow estimated range to X * (0.66 - 2.4) = and our wider range to X * (0.15 - 11). Or, about 1000-4000 for the low-confidence range and 270-19000 for the high-confidence range, with X as the geometric mean. As you see, 1700 indeed turns out to be a good single number to summarize these ranges, much better than was at first obvious. Let's say we don't know how the author has arrived at his range? Taking the geometric mean of these ranges yields about 2000-2200, which is still a pretty good estimate! So regardless of whether the range is simply the extremes of the individual confidence intervals or more carefully derived, the geometric mean ends up being a useful number.

In summary, with some basic statistical and methodological assumptions, the geometric mean ends up being a useful point estimator for a range obtained by multiplying together a number of uncertain factors - and it's pretty good even if we don't know exactly how a range was arrived at by an author, as long as we make the same assumptions about how the underlying distribution looks.

This ended up being a very mathy response, but I hope it gives you the insight you're looking for.