We use cookies, including third-party cookies from Google to serve personalized ads through AdSense, to operate this site and understand how it is used. By continuing to browse, you accept this use. See our Privacy Policy and Terms of Use for details, including how to opt out of personalized advertising.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    chatgpt image jul 21, 2026, 04 34 30 pm
    4 Core Benefits of Predictive Maintenance after Vibration Analysis
    10 Min Read
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results -- AI-generated illustration
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results
    11 Min Read
    chatgpt image jul 13, 2026, 04 23 45 pm
    How Data Analytics Helps Companies Improve User Engagement
    19 Min Read
    chatgpt image jul 13, 2026, 03 59 46 pm
    How Data Analytics Improves Multi-Location Search Strategies
    10 Min Read
    cybersecurity efforts
    How Behavioral Analytics and AI Are Redefining Cybersecurity for Boca Raton Businesses
    14 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: Were the Election Polls Marred by Poor Quality Data?
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Uncategorized > Were the Election Polls Marred by Poor Quality Data?
Uncategorized

Were the Election Polls Marred by Poor Quality Data?

martindoyle
martindoyle
8 Min Read
Were the Election Polls Marred by Poor Quality Data?
Illustration generated with FLUX.2 [klein 4B] via Cloudflare Workers AI.
SHARE

The UK’s general election took place last week, on Thursday 7 May, 2015. It was an election that had been hyped for being ‘too close to call’. According to the polls, the government was likely to be a coalition of one or more, with no party achieving a majority. It could have gone either way.

Contents
  • Precedents in polls
  • How data is sourced
  • Dissecting data collection
  • The consequences of poor quality data

Imagine the shock when the BBC announced the exit poll results: a landslide victory for a single party – the Conservatives.

Election polling companies are reliant on various types of data to come up with accurate predictions. Like any business, they must apply quality control to their data. They must cleanse it, eradicate errors and duplicates, and ensure their contact records are up to date. They need to ensure they don’t call the same person twice, and they must encourage people to give accurate data in response.

How could so many companies get it so badly wrong? And was the data at fault, or was there another gremlin in the machine?

More Read

Big Data: A Brief(ish) History Everyone Should Read
Big Data: A Brief(ish) History Everyone Should Read
The CTOvision.com list of Top Ten CTO Videos
HIPAA Violation Penalties Rise in Response to Data Breaches
Real Business Users and SharePoint
Stupid Analytics Gets Them Talking

Precedents in polls

This is not the first time that polling data has let down the public, politicians and press.

During the US election in 2012, opinion polls predicted a tough campaign for President Obama. He wound up with a comfortable majority. And in 1992, there was a direct comparison with last week’s data disaster. A close race was predicted; the Conservatives won comfortably then, too.

Even this year, in March, polling companies in Israel underestimated support for Netanyahu’s Likud party. He was, in the end, a clear winner.

Even though polls are not binding, they influence voters in the run up to an election, and can even influence the policies that parties formulate as they seek to capture the mood of the electorate. It’s therefore critical that polls can be relied upon. And that makes it a matter of data quality.

How data is sourced

Election polling companies get their data from a variety of sources. YouGov has posted an excellent blog detailing how its surveys work.

In brief, YouGov (and similar organisations) collect data by determining preference for a particular party. Responses are gathered online, and over the phone.

Initially, it’s tempting to think that perhaps different voters have different ways of answering polls. But YouGov adjusts for this already. Labour voters are more likely to use online methods, and Conservative voters the telephone. But it says the data being collected was the same via each method. So we can rule out ‘mode effects’ based on this.

Another culprit is a change in turnout. The polling companies take a small data sample and extrapolate the results, based on the size of the actual sample. So they could be filling in blanks in the wrong way.

Dissecting data collection

YouGov says that there may be “methodological failure” in the way these polls are being conducted.

Interviewers may be asking questions in a loaded way, or influencing the answers by their mere presence.

In the US, polling companies are legally required to manually dial mobile telephone numbers. They do not have to manually dial landlines. This has lead to the landline being favoured, and it’s possible that – culturally – this habit has been persistent in the UK, too. Due to automated calls and marketing campaigns – both relatively new phenomena – some of us view landline telephone calls with suspicion. We want to get off the phone as quickly as possible, so perhaps the data we give is hurried, or we say what we think we should say to get it over with.

They may have asked the wrong questions: how people would vote on local issues, for example, rather than which leader or party was favourable at a national level.

But there may be a more mundane reason. When collecting data, you have to assume that the data is being provided truthfully and accurately. If someone gives you their email address, you need to trust that it will be genuine.

It may be that some people did not supply accurate data to the polling company in the first place: they “said one thing and did another”. This is what Peter Kellner, president of YouGov, thinks is most likely to be the problem. Marketers face this problem all the time: people provide fake information in order to avoid being added to marketing lists. Could it be that people are just less inclined to tell the truth?

The consequences of poor quality data

As providers of data quality software, we have a saying. “Junk in – junk out.”

Whether it’s a company balance sheet, a marketing report or an election poll, the outputs that are generated are only ever going to be as good as the data that was used to build them. This is why data quality is so critical to success and profitability.

Think of it in simple terms. If you compile a list of 1,000 contacts and send a marketing letter to all of them, you need a clean list: one free of mistakes, duplicates, misspellings and ‘gone away’ records.

If you get 500 letters ‘returned to sender’, you can safely assume that your mailing list is seriously decayed. Not only is that inconvenient, but it’s effectively doubled the cost of your campaign: 50% of the effort was wasted.

For polling companies, their entire business model is based on obtaining accurate data, and purifying data to ensure they have a reliable snapshot. They need to take small, frequent samples via polling to gain an understanding of wider trends. This means mistakes are going to be amplified, as they clearly were last week. Now, more than ever, data accuracy is the number one goal.

If telephone polling is less reliable, companies now face a new era, where response rates and accuracy need to be higher. The British Polling Council is conducting an enquiry into the reasons behind last week’s failure. Data quality will undoubtedly be placed under the spotlight. Without it, we may never fully trust election polls again.

Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

Flat editorial illustration: The article's core relationship is the brand protection response workflow: detection of a phishing o
Data & AI Architecture Focus: 6 Best Brand Protection Tools for Phishing and Impersonation
IT Security
Server racks with cloud and user interface panels
Cloud Infrastructure and Workload Migration: A Data-Driven Look at VMware Alternatives in Europe
Cloud Computing Exclusive
Synthetic Data vs Real Web Data: Comparison, Limitations, and Collection Methods  -- AI-generated illustration
Synthetic Data vs Real Web Data: Comparison, Limitations, and Collection Methods 
Big Data Exclusive
Illustration of mobile analytics dashboards with ad performance charts connected to backend databases
11 Best Sisense Alternatives for Embedded Analytics
Business Intelligence Exclusive

Stay Connected

1.2KFollowersLike
33.7KFollowersFollow
222FollowersPin

You Might also Like

Data Mining Interview: Dr. A. Fazel Famili
Uncategorized

Data Mining Interview: Dr. A. Fazel Famili

8 Min Read
6 SMB Technology Trend Predictions for 2016
Uncategorized

6 SMB Technology Trend Predictions for 2016

6 Min Read
The lesson of the Palace of Culture and Science
Uncategorized

The lesson of the Palace of Culture and Science

11 Min Read
How Big Data, and Critical Thinking, Lead to Business Value
Uncategorized

How Big Data, and Critical Thinking, Lead to Business Value

12 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

5 Great Tips for Using Data Analytics for Website UX
5 Great Tips for Using Data Analytics for Website UX
Big Data
Artificial Intelligence for eCommerce: A Closer Look
Artificial Intelligence for eCommerce: A Closer Look
Artificial Intelligence

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-26 SmartData Collective. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?