Cookies help us display personalized product recommendations and ensure you have great shopping experience.

By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    unusual trading activity
    Signal Or Noise? A Decision Tree For Evaluating Unusual Trading Activity
    3 Min Read
    software developer using ai
    How Data Analytics Helps Developers Deliver Better Tech Services
    8 Min Read
    ai for stock trading
    Can Data Analytics Help Investors Outperform Warren Buffett
    9 Min Read
    media monitoring
    Signals In The Noise: Using Media Monitoring To Manage Negative Publicity
    5 Min Read
    data analytics
    How Data Analytics Can Help You Construct A Financial Weather Map
    4 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: Responding to a Follower’s Question: Why Keep Data Replication to a Minimum?
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Big Data > Data Mining > Responding to a Follower’s Question: Why Keep Data Replication to a Minimum?
Data MiningData QualityData Visualization

Responding to a Follower’s Question: Why Keep Data Replication to a Minimum?

Rob Armstrong
Rob Armstrong
4 Min Read
SHARE

I got an e-mail from one of my followers (I say this in the hopes there are many more!).  In my blog post “What will you tolerate?” I provided a sample of guiding principles.  One of them was the suggestion that data replication be kept to a minimum.  The reader wanted to get a bit more depth to that point.

I got an e-mail from one of my followers (I say this in the hopes there are many more!).  In my blog post “What will you tolerate?” I provided a sample of guiding principles.  One of them was the suggestion that data replication be kept to a minimum.  The reader wanted to get a bit more depth to that point.

Going with the theory that if one person in the audience has that question there may be several more with the same thought, I wanted to just clarify and expand on that point.  I will give the shout out to Jon and thank him for the question (as well as correcting some of my typos).

More Read

Privacy Policy Perspectives
DIALOG Sodexo – Workforce Management
Jeff Hawkins: Brain science is about to fundamentally change…
Search Innovation: Why Can’t We All Just Get Along?
First Look – RuleXpress

Editor’s note: Rob Armstrong is an employee of Teradata. Teradata is a sponsor of The Smart Data Collective.

So why do I suggest that data replication be minimized.  There are several reasons beyond the very obvious one of disk storage and cost to maintain.

The main point about this guiding principle is that once the data has been cleansed, transformed, and integrated into the core data warehouse, the access should be against that data directly (or through views).  There is very little reason to then extract the data to another database or platform for analytics.  Many people will extract the data into data marts, excel, or other applications. 

Often times this is justified by claiming performance factors, IT barriers, or a variety of other issues.

Whatever the reason, this duplication of data is a problem and should be avoided when possible.  When data is replicated out very rarely do the data rules, data quality, and auditing trails accompany the extract.  This leads to users taking data, possibly transforming it in their on spreadsheets and then sharing that extract with others.  Now the data in the data warehouse no longer matches any reports or analytics from the extracts.  This lead to confusion and finger pointing about where answers are coming from and who’s answers are correct.  Added to this is a problem when a user decided they want to “drill down” from manipulated data but the underlying data in the warehouse no longer matches the reports.

Now this is not to say there is never a time that replicating data is justified.  Clearly, you will need to replicate data for disaster recovery systems.  You may also want to replicate data into a test environment so new applications can be developed and tested against “real data”.  These cases are reasonable as the data is audited for consistency and do not become the source of new analytics.

You may also have to overcome a real technical issue such as an business critical (with proven value) application that requires data to be co-located with the process.  In this case care should be taken to really document what the technical issue is and how it needs to be resolved.  Finally, there needs to be the understanding that the application will be pointed back to the core warehouse once the issue is resolved.  This of course leads us to the whole role of governance but that is another blog.

That help?

Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

Hidden AI, a risk?
Hidden AI, Real Risk: A Governance Roadmap For Mid-Market Organizations
Artificial Intelligence Exclusive Infographic
unusual trading activity
Signal Or Noise? A Decision Tree For Evaluating Unusual Trading Activity
Analytics Exclusive Infographic
Ai agents
AI Agent Trends Shaping Data-Driven Businesses
Artificial Intelligence Exclusive Infographic
Why Businesses Are Using Data to Rethink Office Operations
Why Businesses Are Using Data to Rethink Office Operations
Big Data Exclusive

Stay Connected

1.2KFollowersLike
33.7KFollowersFollow
222FollowersPin

You Might also Like

Predictive Analytics World

2 Min Read

The Physical Size of Big Data [INFOGRAPHIC]

1 Min Read
Data Clustering
Big DataData ManagementData MiningData Warehousing

Overcoming the Challenges of Big Data Clustering

5 Min Read

Smart meters need smart systems, not a better user interface

6 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

AI and chatbots
Chatbots and SEO: How Can Chatbots Improve Your SEO Ranking?
Artificial Intelligence Chatbots Exclusive
ai is improving the safety of cars
From Bolts to Bots: How AI Is Fortifying the Automotive Industry
Artificial Intelligence

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-25 SmartData Collective. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?