Are public opinion polls from the past reliable? Were they more prone to error back then? Are they reliable in general?

by Dwitt01

I ask this because I’ve seen many public opinion polls from past decades that vary wildly or contradict each other, and others that have results so extreme in a direction that it’s hard to believe.

How have the methods of collecting public opinion data evolved and are they in general reliable?

270-

I'm going to talk exclusively about the history of American opinion polling, especially political opinion polling here, because that's what I know best. Graham Kalton's paper "Developments in Survey Research over the Past 60 Years: A Personal Perspective" (https://onlinelibrary.wiley.com/doi/epdf/10.1111/insr.12287) has a decent overview of the infancy of polling internationally in the late 1800s and early 1900s, if you're further interested in that.

The general underlying principle of polling is that you try to contact a sample that's representative of the general population as possible, and if you can't, you make up for it by weighting individual responses more or less heavily to try and approximate the makeup of the general population.

So the accuracy of a poll depends primarily on how representative a sample you can get (which has waxed and waned over the decades) and how sophisticated a weighting adjustment the pollsters are capable of making (which pollsters have generally become better at, largely because of better data to work with).

In the very early stages of polling, a century ago, it was mostly conducted by newspapers and magazines who would send out poll questions by mail to their subscribers, sometimes expanding that mail to people who had addresses on public record (through phonebooks and similar things), and then just count the responses they got from people who mailed them back.

Needless to say, this method does horrible on basically all fronts-- only a small percentage of people are even theoretically possible to be reached, as magazine subscribers or people who could afford a phone in those days they're not particularly representative of the general population, and the people sending back responses were likely not even representative of that universe, and the magazines did nothing to account for all of that, resulting in blunders like the Literary Digest famously predicting a landslide victory for Alf Landon in 1936.

In the mid-to-late-30s along came George Gallup, the founder of the still existing Gallup polling institute, who understood most of the issues I raised at the top. He used something known as "quota sampling" for his polls. He would set out parameters for a few variables for how many of them his polls were supposed to have-- say, 50% women, 10% African-American, 25% from the South--, then picked randomly selected towns from across the country and sent interviewers there to talk to people according to the quotas he set. Later, those face-to-face interviews would be supplanted more and more by phone polling, but in the late 30s the rate of phone ownership was still quite low.

This was certainly a giant leap over the earlier magazine polling, but by modern standards it still had significant issues-- sampling within the quotas for example was haphazard at best. If you're an interviewer told to talk to ten white men in Galena, Ill., your biases, conscious or unconscious, will probably do a lot to make them unrepresentative of white men in Galena overall. Maybe you just talk to people you perceive as more approachable, who might be more cleanly dressed, wealthier people. For scholars of historical opinion polling, Adam Berinsky and Eric Schickler have developed some statistical reweighting methods to salvage as much use from that data as possible for analysis today, but that wasn't something that was done at the time.

Then in the late 50s/early 60s, landline phone ownership became ubiquitous enough that a lot of pollsters switched from a mix of mail and face-to-face interviews to phone polling, often with early computer systems dialing random phone numbers and connecting operators once a person picked up the phone. This was in some ways a golden age of polling-- nearly everybody was contactable on one single mode of communication, the phone, random digit dialing allowed everyone in that group to be contactable, and people in that time were generally very willing to pick up the phone and talk to pollsters. Response rates north of 75% were not unheard of in that time period. Weighting was still usually fairly simple, but with samples being actually fairly representative of the population, that was acceptable.

Then we come to the late 90s and early 00s, and new problems that polling is still dealing with today emerge. More and more people give up landlines for cell phones. Response rates drop precipitously because people are just inundated with calls from telemarketers, scammers and too many pollsters, so more and more people, especially on cell phones, don't accept calls from unknown numbers anymore, and the people who do still take phone polls are not even remotely representative of the general population anymore-- they skew widely old, just as one easy example, but are also off in many other ways.

A far cry from the 70%+ response rates decades ago, pollsters today can count themselves lucky to get a 5% response rate on a phone poll. As a result, good pollsters have adopted a wide range of other modes of communication to reach people--going back to mail and offering people financial incentives to mail back the surveys, online panels, even SMS polling. That samples now come from a bunch of disparate sources, each representing a different small slice of the population, has caused pollsters to develop more and more sophisticated weighting techniques, aided by commercial files of people with hundreds of variables being available on them that pollsters of the past could only have dreamed of.

When this weighting is done well and captures all the ways in which people polled are meaningfully different from the general population, polling can still do really well, but it's turned into a bit of a race where pollsters are trying to catch on to ever new trends in which people they can't reach-- a common hypothesis for the polling misses in 2016 for example is that people with lower levels of social trust were both less likely to answer polls and less likely to vote for Democrats, not exactly something that is easy to catch or model away.

Finally,

are they in general reliable?

Depends on what standard for reliable you have. Polls generally report a margin of error (for most public polls that's going to be in the 3-5% rate). Given a value from a poll, the probability distribution for the true value ideally is centered around the value from the poll with a normal distribution around it, and the margin of error represents two standard deviations in each direction.

But that presupposes perfect sampling conditions-- that the poll had an equally good chance to get a response from everybody. As we know now, that's never actually the case.

Then on top of that there's several other sources of error: First, the weighting process that makes the poll more representative results in some loss of accuracy--in layman's terms, weighting might from a poll of 500 unrepresentative people "salvage" something that's as good as a poll of 300 representative people would have been. That's called the "design effect", which is multiplied with the margin of error and typically can be in the 1.2-1.5 range for most polls.

Then you have the bias introduced by the imperfections of the sampling that the weighting process could not make up for. This is even more dangerous because it means that the poll will be off in one direction consistently, not randomly in either direction-- something that can't be fixed by just conducting more interviews, and is unknowable ahead of time in an instance like an election where the outcome is verified later, and completely unknowable when it comes to opinion questions where the outcome is never verified. But generally that's still only going to be 2-3% at most in most polling, based on elections where we can verify the accuracy of the polls.

If you want a rule of thumb, I'd say that in polling that is conducted along best practices, you can be very confident that the true value is within 10% of what the poll predicts and have better than even odds that it's within 5%. (This is for a question asked of the entirety of the respondents, not values for subgroups).

Of course, on top of that you have problems with question wording and all that-- the poll at best can only tell you what people think of the exact question wording used in the poll, not what they think of an issue in general. Especially for policy questions where people don't have very strong opinions going in, you can get widely different results based on the exact way the question is asked.