Cookies help us display personalized product recommendations and ensure you have great shopping experience.

By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    New Data Analytics Breakthroughs Give eCommerce Startups a Fighting Chance
    New Data Analytics Breakthroughs Give eCommerce Startups a Fighting Chance
    6 Min Read
    How Data Analytics Is Reshaping Patient Financing Decisions
    How Data Analytics Is Reshaping Patient Financing Decisions
    13 Min Read
    business using business intelligence
    How to Use a Competitive Intelligence Dashboard to Turn Market Data Into Smarter Marketing Decisions 
    9 Min Read
    unusual trading activity
    Signal Or Noise? A Decision Tree For Evaluating Unusual Trading Activity
    3 Min Read
    software developer using ai
    How Data Analytics Helps Developers Deliver Better Tech Services
    8 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: Another Step for Google into Business Analytics? EIM with Google Refine 2.0
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Business Intelligence > Another Step for Google into Business Analytics? EIM with Google Refine 2.0
Business Intelligence

Another Step for Google into Business Analytics? EIM with Google Refine 2.0

Timo Elliott
Timo Elliott
4 Min Read
SHARE

Google has announced Google Refine 2.0, “a power tool for data wranglers”. It’s an open-source tool for cleaning and enhancing “messy” data sets, including cleaning up inconsistencies, transforming them from one format into another, and extending them with new data from external web services or other databases.

More Read

IBM Brings Business Analytics to Apple iPad
Will AI Developments Help Open Banking Take Off?
Generative AI: Unlocking New Revenue Streams for Your Business
How MapR’s M7 Platform Improves NoSQL and Hadoop
Musk vs. Zuckerberg: Who’s Right About AI?

How does it shape up from a corporate point of view? Is this another step for Google into enterprise information management?

The new version includes:

“new extensions architecture, a reconciliation framework for linking records to other databases (like Freebase), and a ton of new transformation commands and expressions.”

The tool runs on your desktop (even though you’re accessing through a browser), so you don’t have to worry about data security.

Here are some videos about the new product. First, an introduction to how you can use the tool to do basic manual cleansing of data:

  • Converting various badly-entered variants of “FFP” into “Firm Fixed Price”
  • Helping “cluster” groups of similar data together using heuristics
  • Using expressions to change distributions using log functions
  • Identifying problems (zero values, errors between millions/billions, etc.)

 

Next, data transformations. The tool makes it easy to convert information in a basic HTML list into a nicely formatted table, using filtering and an expression language:

  • Isolating certain rows
  • Simple conversions, e.g. removing bolding
  • Extracting part of the values to a new column
  • Splitting existing values into new columns

The extractions used then be exported into a standard like JSON, and used to convert similarly-formatted datasets.

 

Finally, you can use Google Refine to augment your data with data from web services and how to link your data with databases such as Freebase:

  • Calling web services to add geo-coding to address information
  • Using Google’s language detection service to identify the language of different values
  • Doing database joins with external data sources (Google calls this “reconciliation”)
  • Freebase has a nice service that will automatically figure out what type of values you have in your data (e.g. movie names), and matches them to appropriate records, and offers you options in case of ambiguity (e.g. there are several movies containing the word “Terminator”)
  • Once you have a match, you can choose from other fields (such as the movie release date, etc.)

 

Overall, it looks like Google Refine 2.0 is great free option for correcting and correlating small, real-world messy data sets from the web — but you have to be a JSON ninja to get the most out of it.

It’s clear that you could already use this in a corporate context — there are many one-off projects that require manual collection of information from different sources. And organizations suffer from just the same types of information fragmentation as the general web.

With a more user-friendly interface and (presumably) more robust scalability, this type of tool could be of great benefit for people trying to pragmatically cobble together  information from various data sources (a few spreadsheets, an external web site, the corporate data warehouse, etc.)

It’s not clear that corporate data is or ever will be a priority for Google — the entire enterprise information management market is a rounding error compared to their advertising empire. In the meantime, anything that can help people get the data they need, cheaply, should be welcomed.

This type of solution will never replace the need for a robust enterprise information platform, but the need for “messy” solutions to answer real-world business questions is a frequently-underestimated need in business analytics deployments, and Google Refine 2.0 looks like a great tool to add to the workbench.

TAGGED:google
Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

Why Every Small Business Should Care About an AI Image Generator
Why Every Small Business Should Care About an AI Image Generator
Artificial Intelligence Exclusive
ai for instagram reel marketing
How AI Is Changing Instagram Reel Marketing
Artificial Intelligence Exclusive Marketing
protecting data in public
The Importance Of Protecting Sensitive Data In Public Services
Big Data Data Management Exclusive
New Data Analytics Breakthroughs Give eCommerce Startups a Fighting Chance
New Data Analytics Breakthroughs Give eCommerce Startups a Fighting Chance
Analytics Big Data Exclusive

Stay Connected

1.2KFollowersLike
33.7KFollowersFollow
222FollowersPin

You Might also Like

Gapminder: Animating the World’s Data

3 Min Read

Yahoo! CEO Marissa Mayer on Data Portabilty

3 Min Read

Google and Transparency

6 Min Read

Google and the Fairness of Search

0 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

ai chatbot
The Art of Conversation: Enhancing Chatbots with Advanced AI Prompts
Chatbots
AI and chatbots
Chatbots and SEO: How Can Chatbots Improve Your SEO Ranking?
Artificial Intelligence Chatbots Exclusive

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-25 SmartData Collective. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?