In my article, “ Data Integration Roadmap to Support Big Data and Analytics,” I detailed a five step process to transition traditional ETL infrastructure to support the future demands on data integration services. It is always helpful if we have an insight into the end state for any journey. More so for the data integration work that is constantly challenged to hit the ground running.
A big data integration ecosystem connects sources, big data storage, data discovery, the enterprise data warehouse, business intelligence, and analytics. Design the connections around how each component uses data: discovery needs quick access, reporting needs dependable historical results, and predictive modeling needs reusable access to structured and unstructured data.
There are two major architectural changes that are shaking the traditional integration platforms warranting a journey into the future state. First, the ability and needs for organizations to store and use big data. Most of the big data has always been available for a longtime, but tools and techniques have expanded the ways to process it for the business benefits. Second, the need for predictive analytics based on the history or patterns of past or hypothetical data driven models. While the business intelligence deals with what has happened, business analytics deal with what is expected to happen. The statistical methods and tools that predict the process outputs in the manufacturing industry have been there for several decades, and these methods also support broader applications of predictive analytics using organizational data assets.
The diagram below depicts the most common end state for the data integration ecosystem. There are six major components in this system.

Sources
The first component is the set of the sources for structured or unstructured data. With the addition of cloud hosted systems and the mobile infrastructure, the size, velocity and complexity of the traditional datasets began to multiply significantly. Growing volumes of structured and unstructured data continue to broaden the demands on integration platforms. With this level of growth, data sources and their sheer volume forms the main component of the new data integration ecosystem. Data integration architecture should enable multiple strategies to access or store this diverse, volatile and exploding amount of data.
Big Data Storage
While the big data storage systems like Hadoop provide good means to store and organize large volumes of data, processing it to extract the snippets of useful information requires a separate processing design. Hadoop is one distributed storage option alongside cloud object storage and data lake patterns. Map/Reduce provided a processing model for large datasets and opened up doors to many new data analytics opportunities. The data integration platform needs to build the structure for big data storage and map out its touch points with the other enterprise data assets.
Data Discovery Platform
The data discovery platform is a set of tools and techniques that work on the big data file system to find patterns and answers to questions business may have. Presently, it is mostly an Adhoc work and organizations still have difficulty putting a process around it. Most people compare the data discovery activity with the gold mining. Only that in this case, by the time one completes mining gold, the silver becomes more valuable. In other words, what is considered valuable information now may be history and unusable only a few hours later. The data integration architecture should encompass this quick and fast paced data crunching enforcing the data quality and the governance. As I detailed in my article, “ Data Analytics Evolution at LinkedIn – Key Takeaways,” strategies such as LinkedIn’s “three second rule,” can drive the data integration infrastructure to be very responsive to meet the end user adaptation needs. According to LinkedIn, the repeated Adhoc requests are systemically met by developing data discovery platform that has a very high degree of reusability of the lessons learned.
Enterprise Data Warehouse
The traditional data warehouses will continue to support the core information needs, but will have to encompass the new features to integrate better with the unstructured data sources and also the performance demands of the analytics platforms. Organizations have begun to develop new approaches to isolate the operational analytics from deep analytics on the history for strategic decisions. The data integration platform should be versatile to isolate the operation information from the strategic longer-term data assets. Also the data integration infrastructure needs to be more temperamental to enable quick access to most widely and frequently accessed data.
Business Intelligence Portfolio
The business intelligence portfolio will continue to focus on the past performance / results even though there would be increased demands for operational reporting and performance. The evolving needs of self-service BI and mobile BI will continue to post architectural challenges to the data integration platforms. One other critical aspect would be BI portfolio’s ability to integrate with the data analytics portfolio. This need may further increase the demands on enterprise information integration.
Data Analytics Portfolio
There is a reason why they call people working with data analytics as data scientists. Analytical work that goes on within this portfolio need to deal with business as well as data problems and the data scientists need to work their way through building the predictive models that add value to the organization. Data integration platform plays two roles to support the analytics portfolio. First, data integration ecosystem should enable access to structured or unstructured data for analytics. Second, enable re-usability of the past analytics activity to make the field more of an engineering activity than science by reducing the scenarios requiring reinventing the wheel.
How the six components map to the layers of a data platform
Viewed as layers of a data platform, the six components group by their roles. Sources provide the structured and unstructured data. Big Data Storage stores and organizes large volumes, while the Enterprise Data Warehouse supports core information needs and separates operational information from longer-term historical assets. The Data Discovery Platform works on stored big data to find patterns and answer business questions, requiring fast access alongside data quality and governance.
The Business Intelligence Portfolio and Data Analytics Portfolio use data for different purposes: BI focuses on past performance and operational reporting, while analytics builds predictive models using structured and unstructured data and reuses past analytical work. Data integration connects these roles, mapping storage touch points to other enterprise assets and supporting the different access and performance demands of discovery, reporting, and analytics.
The data integration ecosystem encompasses processing very large volumes of data and deals with very diverse demands to work with many varieties of sources of data as well as the end user base.
Frequently Asked Questions
What is an ecosystem in big data?
A big data ecosystem is the connected set of data sources, storage systems, processing tools, and analytical applications used to turn large, diverse datasets into useful information. In this article, it includes six components: Sources, Big Data Storage, Data Discovery Platform, Enterprise Data Warehouse, Business Intelligence Portfolio, and Data Analytics Portfolio.
Can you give me an example of a data ecosystem?
The six-component end state described in this article is an example. Sources supply structured and unstructured data to Big Data Storage. A Data Discovery Platform works on the stored big data to find patterns and answers to business questions. The Enterprise Data Warehouse keeps the historical information behind core reporting, the Business Intelligence Portfolio reports on past performance, and the Data Analytics Portfolio builds predictive models. The data integration platform connects all six.
What are the 5 layers of a data platform?
This article does not define five layers. It describes six components: Sources, Big Data Storage, Data Discovery Platform, Enterprise Data Warehouse, Business Intelligence Portfolio, and Data Analytics Portfolio. Read as platform layers, sources bring the data in, storage and the warehouse hold it, discovery explores it, and business intelligence and analytics put it to use.


