By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    data Analytics instagram stories
    Data Analytics Helps Marketers Make the Most of Instagram Stories
    15 Min Read
    analyst,women,looking,at,kpi,data,on,computer,screen
    What to Know Before Recruiting an Analyst to Handle Company Data
    6 Min Read
    AI analytics
    AI-Based Analytics Are Changing the Future of Credit Cards
    6 Min Read
    data overload showing data analytics
    How Does Next-Gen SIEM Prevent Data Overload For Security Analysts?
    8 Min Read
    hire a marketing agency with a background in data analytics
    5 Reasons to Hire a Marketing Agency that Knows Data Analytics
    7 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-23 SmartData Collective. All Rights Reserved.
Reading: Tweety Bird and Aha! Moments
Share
Notification Show More
Aa
SmartData CollectiveSmartData Collective
Aa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Inside Companies > Tweety Bird and Aha! Moments
Inside Companies

Tweety Bird and Aha! Moments

MIKE20
Last updated: 2010/12/28 at 1:59 PM
MIKE20
4 Min Read
SHARE



About three months ago, I started a data management and ETL project for a pretty big bank. Time was of the essence and the bank brought me in because I can get results. In this post, I explain why an overemphasis on results can be a really bad thing–and why all matching isn’t created equal.When my client advised me of the number of disparate extracts of financial data I would receive, I quickly looked for potential commonalities. As most would do on this type of project, I started with the obvious: GL account. While not unique in most systems I have seen, they are often at least part of a multi-field index. In other words, I can almost always sum debits and credits by GL account 1000, for example, especially when I bring in a field like company.

Dealing with Suboptimal Data and Other Limitations

Unfortunately, a few of the extracts only contained account descriptions. When I explained this limitation to my boss, we had the following exchange:

Boss: Well, why not join on description instead of account number?

More Read

analyzing big data for its quality and value

Use this Strategic Approach to Maximize Your Data’s Value

7 Data Lineage Tool Tips For Preventing Human Error in Data Processing
Preserving Data Quality is Critical for Leveraging Analytics with Amazon PPC
Quality Control Tips for Data Collection with Drone Surveying
3 Huge Reasons that Data Integrity is Absolutely Essential

Me: It’s possible, but I never recommend it. In many systems, these descriptions are not standardized. One missing character, misplaced space, or mutant letter will result in missing data downstream and other problems that are worth trying to address from the beginning.

Boss: We have no other choice. IT won’t change the extract from XYZ financial system. Gotta make do….

Me (in Tweety Bird voice): OK…but I got a baaaaaad feeling about this.

Fast forward to the end of the project. Our reports were off–way off. What’s more, this wasn’t a fundamental design issue. So far off that I had to rejigger a bunch of previously working queries and validation routines. This took a few days and caused other problems. We had to break things that previously worked.

Sometimes we don’t have a choice in life and in golf. We hit the ball into the woods, can’t find it, have to take a penalty stroke, and walk away with a snowman on a par 3. It happens.

Although my client wasn’t happy that we had so many problems, he fully understood what I told him on day one: “Try not to join on descriptions.” Of course, the bell didn’t ring quite as loud when I initially said that. After he saw the results first-hand, that bell was pretty loud–and wouldn’t stop.

Understand that there are limitations with certain types of data matching, as many on this blog has pointed. Some types are much better than others. Some are too restrictive; others are far too liberal. Think about it.

As one of my friends put so nicely, you don’t go to pick up your kid at nursery school and just grab any eight-year old brunette girl. You probably want you own kid.

Data’s the same way.

What say you?

Read more at MIKE2.0: The Open Source Standard for Information Management

TAGGED: data quality
MIKE20 December 28, 2010
Share This Article
Facebook Twitter Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

smart home data
7 Mind-Blowing Ways Smart Homes Use Data to Save Your Money
Big Data
ai low code frameworks
AI Can Help Accelerate Development with Low-Code Frameworks
Artificial Intelligence
data Analytics instagram stories
Data Analytics Helps Marketers Make the Most of Instagram Stories
Analytics
data breaches
How Hospital Security Breaches Devastate Local Communities
Policy and Governance

Stay Connected

1.2k Followers Like
33.7k Followers Follow
222 Followers Pin

You Might also Like

analyzing big data for its quality and value
Big Data

Use this Strategic Approach to Maximize Your Data’s Value

6 Min Read
data lineage tool
Big Data

7 Data Lineage Tool Tips For Preventing Human Error in Data Processing

6 Min Read
data quality and role of analytics
Data Quality

Preserving Data Quality is Critical for Leveraging Analytics with Amazon PPC

8 Min Read
data collection with drone use
Data Collection

Quality Control Tips for Data Collection with Drone Surveying

9 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

data-driven web design
5 Great Tips for Using Data Analytics for Website UX
Big Data
AI and chatbots
Chatbots and SEO: How Can Chatbots Improve Your SEO Ranking?
Artificial Intelligence Chatbots Exclusive

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
Go to mobile version
Welcome Back!

Sign in to your account

Lost your password?