We use cookies, including third-party cookies from Google to serve personalized ads through AdSense, to operate this site and understand how it is used. By continuing to browse, you accept this use. See our Privacy Policy and Terms of Use for details, including how to opt out of personalized advertising.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    chatgpt image jul 21, 2026, 04 34 30 pm
    4 Core Benefits of Predictive Maintenance after Vibration Analysis
    10 Min Read
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results -- AI-generated illustration
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results
    11 Min Read
    chatgpt image jul 13, 2026, 04 23 45 pm
    How Data Analytics Helps Companies Improve User Engagement
    19 Min Read
    chatgpt image jul 13, 2026, 03 59 46 pm
    How Data Analytics Improves Multi-Location Search Strategies
    10 Min Read
    cybersecurity efforts
    How Behavioral Analytics and AI Are Redefining Cybersecurity for Boca Raton Businesses
    14 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: Virtualization comes of age… again!
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Big Data > Data Warehousing > Virtualization comes of age… again!
Data Warehousing

Virtualization comes of age… again!

Barry Devlin
Barry Devlin
6 Min Read
Virtualization comes of age... again!
Illustrative image generated with OpenAI gpt-image-1.
SHARE

Way back in 1991, when IBM announced the Information Warehouse Framework, one aspect of the content came as a shock to most people who were promoting data warehousing then.  (There were not too many of us at the time to be shocked… as far as I know, no one had yet claimed paternity of data warehousing, and the first popular book on the topic was still a year away.)  The shock was that the announcement included the concept of access to heterogeneous data, to be supported through an alliance with Information Builders Inc., using their product EDA/SQL.  The accepted wisdom in data warehousing at the time and for many years since was that heterogeneous data must be cleansed, reconciled and loaded into the warehouse via ETL tooling and accessed from there.  Information Warehouse was not a great success for IBM, and access to heterogeneous data largely faded from awareness among data warehousing professionals.  That made a lot of sense back then.  Heterogeneous data was very heterogeneous, very complex and very susceptible to performance problems when accessed in an unplanned manner.  “Leave it alone!” was the sensible advice.

Ten years later, I was again faced with the concept of heterogeneous data access as IBM began the work that was later announced as IBM DB2 Information Integrator.  This time, the starting point was federated access to data, initially across relational systems, but with a clear direction to include all types of data.  Again, the market wasn’t really ready for the concept, although there was a wider degree of acceptance of the idea and a number of early adopters began to experiment seriously with implementation. Most data warehousing experts still shook their heads in disbelief…

Fast forward another ten years and access to heterogeneous data is back on the agenda big time, this time under the name virtualization and the launch of Composite 6 yesterday ups the ante again.  (It may be of interest to note that Composite was founded in 2002, right in the middle of the last wave of interest.)  While there is lots of fascinating stuff in the release about improving performance, caching and governance, my attention was drawn particularly to the inclusion of “big data” integration support.  And my concern was how Composite could understand and reliably use the variety of data types, elements and so on, which are typically present in Hadoop files.

My contention is that over the years since 1991, heterogeneous data sources have, generally speaking, become better defined, less complex in terms of structure and content, more easily accessed, and less prevalent.  Until the advent of big data, that is.  In data management terms, big data is like a giant step backwards to the Wild West from modern suburbia: schema–why bother? Metadata–who needs it, it will be out of date in a day?  Governance–programmers can handle it!

More Read

Connecting the Clouds: Netsuite anounces connectivity to Salesforce.com
Connecting the Clouds: Netsuite anounces connectivity to Salesforce.com
Analyzing Healthcare in Sweden
The Problem with the Relational Database
The Big Data in Teradata
How is Big Data Stored and Managed?

But when I put the question to Dave Besemer, CTO of Composite, the answer I got proved very enlightening.  Not just about Composite’s approach but also about what is going on, perhaps somewhat by stealth, in the world of big data.  Basically, Dave said that Composite accesses big data only via Hive, which provides the basic structural metadata required for virtualization.  And Hive?  Well, Hive defines itself on its own website as, wait for it: “…a data warehouse system for Hadoop that facilitates easy data summarization, ad-hoc queries, and the analysis of large datasets stored in Hadoop compatible file systems. Hive provides a mechanism to project structure onto this data and query the data using a SQL-like language called HiveQL…”

So, as Robert Browning wrote “God’s in his heaven, all’s right with the world” if you are a data management fan.  The big data folks do recognize the value of data management (Hive has been around since 2009) despite some of the NoSQL hype that still continues to turn up in the press.  That’s not to say that Hive needs to be put in front of every set of Hadoop files.  There’s a whole world of distributed Hadoop data that is so transient and/or so specialized that the only sensible way to use it is via a programmatic interface.  But, Composite isn’t going after that stuff; they are focusing on the better defined and managed segment of big data.  And that makes perfect sense.

But there is still a question in my mind that the broader IT community needs to answer:  How are we going to manage and handle the other, much larger segment of big data?  Pat Helland’s article “If You Have Too Much Data, then ‘Good Enough’ Is Good Enough” in the ACM Journal provides some food for thought.

Oh, and by the way, there are a few data warehouse eminences grises who still proclaim that virtualization is evil and that all data has to go through the data warehouse…  Perhaps they’re waiting for the fourth wave?

Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

How Local Service Businesses Can Map Which Neighborhoods Generate the Most Revenue -- AI-generated illustration
How Local Service Businesses Can Map Which Neighborhoods Generate the Most Revenue
Business Intelligence
Best VMware Alternatives in Thailand for Private Cloud and HCI Deployments -- AI-generated illustration
Best VMware Alternatives in Thailand for Private Cloud and HCI Deployments
Cloud Computing Exclusive IT
Using Safety Metrics and Incident Data to Reduce Construction Risk and Insurance Costs -- AI-generated illustration
Using Safety Metrics and Incident Data to Reduce Construction Risk and Insurance Costs
Big Data Exclusive
How Business Intelligence Can Help Small Businesses Build Better Decision Rules -- AI-generated illustration
How Business Intelligence Can Help Small Businesses Build Better Decision Rules
Business Intelligence Business Rules Exclusive

Stay Connected

1.2KFollowersLike
33.7KFollowersFollow
222FollowersPin

You Might also Like

Big Data Success Stories: Take Them with a Grain of Salt

4 Min Read
Why Use Reporting Repositories for Business Intelligence?
Business IntelligenceData Warehousing

Why Use Reporting Repositories for Business Intelligence?

6 Min Read
Why The Last Decade of BI Best-Practice Architecture is Rapidly Becoming Obsolete
AnalyticsBusiness IntelligenceData MiningData WarehousingMarket Research

Why The Last Decade of BI Best-Practice Architecture is Rapidly Becoming Obsolete

25 Min Read
Reference Domains Part III: Collecting Classifications
AnalyticsBest PracticesData Warehousing

Reference Domains Part III: Collecting Classifications

8 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

Chatbots and SEO: How Can Chatbots Improve Your SEO Ranking?
Chatbots and SEO: How Can Chatbots Improve Your SEO Ranking?
Artificial Intelligence Chatbots Exclusive
From Bolts to Bots: How AI Is Fortifying the Automotive Industry
From Bolts to Bots: How AI Is Fortifying the Automotive Industry
Artificial Intelligence

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-26 SmartData Collective. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?