We use cookies, including third-party cookies from Google to serve personalized ads through AdSense, to operate this site and understand how it is used. By continuing to browse, you accept this use. See our Privacy Policy and Terms of Use for details, including how to opt out of personalized advertising.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    What Kind of Problem-Solving Distinguishes Data Analysts From Software Engineers -- AI-generated illustration
    What Kind of Problem-Solving Distinguishes Data Analysts From Software Engineers
    7 Min Read
    chatgpt image jul 21, 2026, 04 34 30 pm
    4 Core Benefits of Predictive Maintenance after Vibration Analysis
    10 Min Read
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results -- AI-generated illustration
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results
    11 Min Read
    chatgpt image jul 13, 2026, 04 23 45 pm
    How Data Analytics Helps Companies Improve User Engagement
    19 Min Read
    chatgpt image jul 13, 2026, 03 59 46 pm
    How Data Analytics Improves Multi-Location Search Strategies
    10 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: Why normalization matters with K-Means
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Big Data > Data Mining > Why normalization matters with K-Means
Data MiningPredictive Analytics

Why normalization matters with K-Means

DeanAbbott
DeanAbbott
4 Min Read
Why normalization matters with K-Means
Illustrative image generated with OpenAI gpt-image-1.
SHARE

A question about K-means clustering in Clementine was posted here. I thought I knew the answer, but took the opportunity to prove it to myself.

I took the KDD-Cup 98 data and just looked at four fields: Age, NumChild, TARGET_D (the amount the recaptured lapsed donors gave) and LASTGIFT. I took only four to make the problem simpler, and chose variables that had relatively large differences in mean values (where normalization might matter). Also, another problem with the two monetary variables is that they are both skewed positively (severely so).

The following image shows the results of two clustering runs: the first with raw data, the second with normalized data using the Clementine K-Means algorithm. The normalization consisted of log transforms (for TARGET_D and LASTGIFT) and z-scores for all (the log transformed fields, AGE and NUMCHILD). I used the default of 5 clusters.

Here are the results in tabular form. Note that I’m reporting unnormalized values for the “normalized” clusters even though the actual clusters were formed by the normalized values. This is purely for comparative purposes.

More Read

Rexer Analytics Survey – are data miners too focused on their models?
Rexer Analytics Survey – are data miners too focused on their models?
Interview KXEN Bruno Delahaye
So What is Business Analytics & Its Various Components?
Business Analytics – Opposition or Proposition?
NCAA Bracketology and Other Sports Analytics Winners

Note that:
1) the results are different, as measure by counts in each cluster
2) the unnormalized clusters are dominated by TARGET_D and LASTGIFT–one cluster contains the large values and the remaining have little variance.
3) AGE and NUMCHILD have some similar breakouts (40s with more children and 40s with fewer children for example).

So, the conclusion is (to answer the original question) K-Means in Clementine does not normalize the data. Since Euclidean distance is used, the clusters will be influenced strongly by the magnitudes of the variables, especially by outliers. Normalizing removes this bias. However, whether or not one desires this removal of bias depends on what one wants to find: sometimes if one would want a variable to influence the clusters more, one could manipulate the clusters precisely in this way, by increasing the relative magnitude of these fields.

One last issue that I didn’t explore here, is the effects of correlated variables (LASTGIFT and TARGET_D to some degree here). It seems to me that correlated variables will artificially bias the clusters toward natural groupings of those variables, though I have never proved the extent of this bias in a controlled way (maybe someone can point to a paper that shows this clearly).

Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

Technical architecture diagram and decision framework for ai-powered logo generation: design.com vs looka
AI-Powered Logo Generation: Design.com vs Looka
Artificial Intelligence Exclusive
Retailers Should Stop Treating Every Stockout as Equal -- AI-generated illustration
Retailers Should Stop Treating Every Stockout as Equal
Business Intelligence Exclusive
Best Vibe Coding Cleanup Specialists in the USA: Fix or Rebuild? -- AI-generated illustration
Best Vibe Coding Cleanup Specialists in the USA: Fix or Rebuild?
Development Exclusive
Using Warehouse, Transportation and Order Data to Plan Distribution-Center Capacity -- AI-generated illustration
Using Warehouse, Transportation and Order Data to Plan Distribution-Center Capacity
Big Data Exclusive

Stay Connected

1.2KFollowersLike
33.7KFollowersFollow
222FollowersPin

You Might also Like

Geeks are Chic
Data Mining

Geeks are Chic

5 Min Read
From Master Data to Master Graph
AnalyticsBig DataBusiness IntelligenceCRMData ManagementData MiningData VisualizationMarketingMarketing AutomationModelingPredictive Analytics

From Master Data to Master Graph

16 Min Read
Is Data Analytics Ushering in the Modern Age of Weather Forecasting?
Analytics

Is Data Analytics Ushering in the Modern Age of Weather Forecasting?

6 Min Read
We Need an "Internet of Not Only Customers"
AnalyticsBest PracticesBig DataBusiness IntelligenceCloud ComputingCRMData ManagementData MiningData QualityData VisualizationData WarehousingMarketingModelingSoftwareSQL

We Need an “Internet of Not Only Customers”

15 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

How To Get An Award Winning Giveaway Bot
How To Get An Award Winning Giveaway Bot
Big Data Chatbots Exclusive
Artificial Intelligence for eCommerce: A Closer Look
Artificial Intelligence for eCommerce: A Closer Look
Artificial Intelligence

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-26 SmartData Collective. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?