We use cookies, including third-party cookies from Google to serve personalized ads through AdSense, to operate this site and understand how it is used. By continuing to browse, you accept this use. See our Privacy Policy and Terms of Use for details, including how to opt out of personalized advertising.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    chatgpt image jul 21, 2026, 04 34 30 pm
    4 Core Benefits of Predictive Maintenance after Vibration Analysis
    10 Min Read
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results -- AI-generated illustration
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results
    11 Min Read
    chatgpt image jul 13, 2026, 04 23 45 pm
    How Data Analytics Helps Companies Improve User Engagement
    19 Min Read
    chatgpt image jul 13, 2026, 03 59 46 pm
    How Data Analytics Improves Multi-Location Search Strategies
    10 Min Read
    cybersecurity efforts
    How Behavioral Analytics and AI Are Redefining Cybersecurity for Boca Raton Businesses
    14 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: Big Data Blasphemy: Why Sample?
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Data Management > Best Practices > Big Data Blasphemy: Why Sample?
AnalyticsBest PracticesData MiningData QualityPredictive AnalyticsStatisticsWeb Analytics

Big Data Blasphemy: Why Sample?

metabrown
metabrown
8 Min Read
Big Data Blasphemy: Why Sample?
Photo by Godfrey Atima on Pexels (https://www.pexels.com/photo/close-up-of-codes-on-a-computer-screen-4976712/)
SHARE

Since data mining began to take hold in the late nineties, “sampling” has become a dirty word in some circles. The Big Data frenzy is compounding this view, leading many to conclude that size equates to predictive power and value. The more data the better, the biggest analysis is the bestest.

Except when it isn’t, which is most of the time.

Data miners have some legitimate reasons for resisting sampling. For starters, the vision of data mining pioneers was empowerment of people who had business knowledge, but not statistical knowledge, to find valuable patterns in their own data and put that information into use. So the intended users of data mining tools are not trained in sampling techniques. Some view sampling as a process step that could be omitted, provided that the data mining tool can run really, really fast. Current data mining tools make sampling quite easy, so this line of resistance has withered quite a lot.

The other significant reason why data miners often choose to use all the data they have, even when they have quite a lot, is that they are looking for extreme cases. They are on a quest in search of the odd and unusual.  Not everyone has pressing needs for this, but for those who have, it makes sense to work with a lot of data. For example, in intelligence or security applications, only a few cases out of millions may exhibit behavior indicative of threatening activity.  So analysts in those fields have a darned good reason to go whole hog.

More Read

Worst Practices While Deploying a Predictive Model
Worst Practices While Deploying a Predictive Model
Big Data and Analytics Must Be Properly Matched
Big Data Solutions in the AWS Platform
Electronic Substitution in the New Economy
Data Analytics Helps Beginning Forex Traders But Doesn’t Replace Common Sense

It’s mighty odd, though, that many people who have no clear business reason for obsessing over rare cases get their panties wound up in a bunch at the mere mention of sampling. The more that I talk to these outraged investigators, the more I believe that this simply reflects poor grounding in data analysis methods.

To put it bluntly, if you don’t sample, if you don’t trust sampling, if you insist that sampling obscures the really valuable insights, you don’t know your stuff. The best analysts, whether they call themselves analysts, scientists, statisticians or some other name, use sampling routinely. But there are many “gurus” out there spreading misleading information. Don’t buy what they are selling.

So what is a sample? A sample is small quantity of data.

Small is relative. A poll to predict election outcomes could get by with no more than a couple of thousand respondents, perhaps just a few hundred, to gauge the attitudes of millions of voters. A vial of your blood is sufficient to assess the status of all the blood in your body. Even a massive data source, with millions of millions of rows of data is still just a sample of the data that could potentially be collected from the big, wide world.

How can you know how big a sample you need? Classical statistics has methods for that, you can learn them. Data mining is much less formal, but the gist would be that if what you discover from your sample still holds water when you test on additional data and in the field, it was good enough.

In data analysis, we select samples that are representative of some bigger body of data that interests us. The big body of data does not refer to the data in your repository. In statistical theory, it’s called the “population,” which is more of an idea than a thing. The population means all the cases you want to draw conclusions about. So that may include all the data in your repository, as well as data that has been recorded in some other resources you cannot access. It can also include cases that have taken place, but for which no data was recorded, and even cases which have not yet occurred.

You may have heard the term “random sample.” This means that every case in the population has an equal opportunity to get in the sample.  The most fundamental assumption of all statistical analysis is that samples are random (ok, there are variations on that theme, but we’ll save that for another day). In practice, our samples are not perfectly random, but we do our best.

If you use all the data in your Big Data resource, you’re not really avoiding sampling. No doubt you will use your analysis to draw conclusions about future cases – cases that are not in your resource today. So your Big Data is still just a very, very big sample from the population that matters to you.

But, if you have it, why not use it? Why wouldn’t you use all the data available?

More isn’t necessarily better. Analyzing massive quantities of data consumes a lot of resources, in computing power, storage space, in the patience of the analyst. Assuming that the resources are even available, the clock is till ticking, and every minute you are waiting for meaningful analysis is a minute when you don’t have the benefit of information that could be put to use in your business. The resources used for just one analysis involving a massive quantity of data could be sufficient to produce many useful studies if you’d only use smaller samples.

Resources are not the only issue. There’s also the little matter of data quality. Is every case in your repository nice and clean? Are you sure? What makes you sure? How about investigating some of that data very carefully and looking for signs of trouble? Much easier to assure yourself that a modest-sized sample is nicely cleaned up than a whole, whopping repository. Data quality is a whole lot more valuable than data quantity.

You see, ladies and gentlemen of the analytic community, that sampling is not a dirty word. Sampling is a necessary and desirable item in the data analysis toolkit, no matter what type of analysis you require. If you’re not familiar or comfortable with it, change your ways now.

©2012 Meta S. Brown

TAGGED:big datadata samplingrandom sampling
Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

Flat editorial illustration: The article's core relationship is the alignment between customer behavioral data (visit frequency,
Data-Driven Loyalty: How Restaurants Use Behavioral Analytics to Optimize Revenue
Exclusive
Flat editorial illustration: The article's core relationship is that reliable eCommerce attribution depends on a unified, well-st
How eCommerce Data Teams Can Build Attribution That Holds Up
Big Data Exclusive
Flat editorial illustration: The article's core relationship is the contrast between fragmented inherited data infrastructure (wh
Data Stack Consolidation as a Data Quality and Governance Strategy for Mid-Market Teams
Big Data Exclusive
Emergency responder and nurse reviewing tablet with data dashboards
Evaluating Workforce Assessment Tools: Looking Beneath the Dashboard at Psychometric Data
Exclusive Software

Stay Connected

1.2KFollowersLike
33.7KFollowersFollow
222FollowersPin

You Might also Like

Big Data Insights Drive Surge In Digital Marketing ROI
Big DataExclusiveMarketing

Big Data Insights Drive Surge In Digital Marketing ROI

6 Min Read
Not Only SQL, Not Only Big Data
Business IntelligenceCommentaryData QualityKnowledge ManagementPolicy and Governance

Not Only SQL, Not Only Big Data

6 Min Read
Data Scalability Leads To New Evolutions In Smart Technology
Big DataExclusive

Data Scalability Leads To New Evolutions In Smart Technology

5 Min Read
How Big Data Has Changed the Financial Industry
Big Data

How Big Data Has Changed the Financial Industry

6 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

Artificial Intelligence for eCommerce: A Closer Look
Artificial Intelligence for eCommerce: A Closer Look
Artificial Intelligence
Chatbots and SEO: How Can Chatbots Improve Your SEO Ranking?
Chatbots and SEO: How Can Chatbots Improve Your SEO Ranking?
Artificial Intelligence Chatbots Exclusive

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-26 SmartData Collective. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?