Cookies help us display personalized product recommendations and ensure you have great shopping experience.

By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    big data analytics in transporation
    Turning Data Into Decisions: How Analytics Improves Transportation Strategy
    3 Min Read
    sales and data analytics
    How Data Analytics Improves Lead Management and Sales Results
    9 Min Read
    data analytics and truck accident claims
    How Data Analytics Reduces Truck Accidents and Speeds Up Claims
    7 Min Read
    predictive analytics for interior designers
    Interior Designers Boost Profits with Predictive Analytics
    8 Min Read
    image fx (67)
    Improving LinkedIn Ad Strategies with Data Analytics
    9 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: Key Words Through Graph Entropy Hierarchical Clustering
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Analytics > Text Analytics > Key Words Through Graph Entropy Hierarchical Clustering
Text Analytics

Key Words Through Graph Entropy Hierarchical Clustering

cristian mesiano
cristian mesiano
4 Min Read
SHARE

In the last post I showed how to extract key words from a text through a principle called graph entropy.
Today I’m going to show another application of the graph entropy in order to extract clusters of key words.

Why
The key words of a document depict the main topic of the content, but if the document is big, often, there are many different sub topics related to the main.

In this perspective, a clusters of keywords should make easier for the reader the identification of the key points of a document.

In the last post I showed how to extract key words from a text through a principle called graph entropy.
Today I’m going to show another application of the graph entropy in order to extract clusters of key words.

More Read

Three Primary Analytics Lessons Learned from 9/11
3 Major Reasons VPN Can Improve Data Security
NGMR Guru Interview with Jeff Jonas of IBM
Can Real-Time Social Analytics Provide Early Indications of Business Results
Uncertainty Coefficients for Features Reduction – Comparison with LDA Technique

Why
The key words of a document depict the main topic of the content, but if the document is big, often, there are many different sub topics related to the main.

In this perspective, a clusters of keywords should make easier for the reader the identification of the key points of a document.

Moreover, imagine to implement a search engine based on clusters of relevant words instead of the common indexing of atomic words: it enables documents comparison, taxonomies definition, and much more!

How
The definition of graph entropy I’m studying on, assigns to each word of the document a relevance score and a sub graph of words topologically closed to it.

The clustering should maximize the relevance score obtained merging two words in the same cluster.

It’s easy to understand that we have to face a combinatoric maximization problem.

The idea is to take advantage of the Simulated annealing (a bit revisited and adapted to the scope) in order to identify sub-optimal merging solution at each step of the merging phase of the hierarchical clustering.

Experiment
I decided to adopt as document test the complete version of the file we used in the last post: Nuclear_weapon.
Here you are the clusters of first 100 relevant words extracted:

The three clusters obtained.
 

It’s interesting to highlight the following considerations:

  • The first cluster merged together words as “material,uranium, plutonium, isotope” and “war, attack, arm“, and also “proliferation, movement, control, development“.
  • The second cluster (which has the lowest rank) aggregates words as “japan, japanese, place, israel, iraq,american“, and “ton, tnt, yeld”  
  • The third cluster (which has the highest rank) describes quite well the primary topic, merging all the most important words of the document! 

Of course, the procedure is still in “incubator” phase, and the accuracy of the clusters rests on the performance of the Annealing clustering (…maybe different algorithms in this context perform better… but just to show a rough solution I guess it’s enough :D)

This is the optimization process for the last merging stage (I presume that temperature schedule requires an adjustment):

Optimization curve through Simulated Annealing Hierarchical Clustering (last merging stage)


Next steps:
Looking forward to receive comments, and suggestions.
…It would be interesting using such methodology to create a new kind of full text search engine, totally independent by frequency of the words and frequency of visits.

The doc
here you are the document parsed and colored through the clustering assignment (have been highlighted just the first 100 relevant features ranked through the Graph Entropy method).
Stay tuned
cristian.


Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

big data analytics in transporation
Turning Data Into Decisions: How Analytics Improves Transportation Strategy
Analytics Big Data Exclusive
AI and fund manager software
AI And The Acceleration Of Information Flows From Fund Managers To Investors
Artificial Intelligence Exclusive
sales and data analytics
How Data Analytics Improves Lead Management and Sales Results
Analytics Big Data Exclusive
ai in marketing
How AI and Smart Platforms Improve Email Marketing
Artificial Intelligence Exclusive Marketing

Stay Connected

1.2kFollowersLike
33.7kFollowersFollow
222FollowersPin

You Might also Like

AnalyticsBig DataBusiness IntelligenceSentiment AnalyticsSocial Media AnalyticsText AnalyticsUnstructured Data

Dark Data: The hidden billion dollar opportunity

4 Min Read
Hadoop in retail
AnalyticsBig DataData VisualizationHadoopMapReduceMarketing AutomationModelingPredictive AnalyticsSentiment AnalyticsSocial DataSocial Media AnalyticsSoftwareSQLText AnalyticsUnstructured DataWeb Analytics

5 Common Use Cases for Hadoop in Retail

5 Min Read
Image
AnalyticsBig DataBusiness IntelligenceData MiningData WarehousingInside CompaniesModelingPolicy and GovernancePredictive AnalyticsPrivacySentiment AnalyticsSocial Media AnalyticsText AnalyticsUnstructured DataWeb Analytics

Facebook’s Big Data: Equal Parts Exciting and Terrifying?

8 Min Read

Taming the Social Media Beast

5 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

giveaway chatbots
How To Get An Award Winning Giveaway Bot
Big Data Chatbots Exclusive
AI chatbots
AI Chatbots Can Help Retailers Convert Live Broadcast Viewers into Sales!
Chatbots

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-25 SmartData Collective. All Rights Reserved.
Go to mobile version
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?