Cookies help us display personalized product recommendations and ensure you have great shopping experience.

By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    image fx (67)
    Improving LinkedIn Ad Strategies with Data Analytics
    9 Min Read
    big data and remote work
    Data Helps Speech-Language Pathologists Deliver Better Results
    6 Min Read
    data driven insights
    How Data-Driven Insights Are Addressing Gaps in Patient Communication and Equity
    8 Min Read
    pexels pavel danilyuk 8112119
    Data Analytics Is Revolutionizing Medical Credentialing
    8 Min Read
    data and seo
    Maximize SEO Success with Powerful Data Analytics Insights
    8 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: A computer program predicts Viral Tweets
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Big Data > Data Mining > A computer program predicts Viral Tweets
Data Mining

A computer program predicts Viral Tweets

ThemosKalafatis
ThemosKalafatis
6 Min Read
SHARE

In the previous post we have seen that the author of a Tweet is the most important factor for making a viral Tweet. This time we will use text mining to score Tweets and see how much viral they could become. Each Tweet is fed to a computer program (an algorithm) and the algorithm responds with the probability each Tweet has to become viral (we assume that when a Tweet receives more than 30 Re-Tweets it is considered viral).

The information that is given to the algorithm is the text of the Tweet and its author. Many other parameters can be taken into consideration such as the time that the Tweet has been posted, the type of the Tweet (e.g., politics, technology, health) or even whether this Tweet is part of a novel subject. Here is the output of the software that performs the predictions:


The number of Re-Tweets is shown in squares. Pay also close attention to the circled text shown above. For each Tweet the most probable outcome is given (‘t’= Tweet will become viral, ‘f’=otherwise) and a confidence for each prediction is given as a number from 0 to 1. As an example, the first Tweet shown above was posted from Paula Abdul saying that she will not return to American Idol. The …


In the previous post we have seen that the author of a Tweet is the most important factor for making a viral Tweet. This time we will use text mining to score Tweets and see how much viral they could become. Each Tweet is fed to a computer program (an algorithm) and the algorithm responds with the probability each Tweet has to become viral (we assume that when a Tweet receives more than 30 Re-Tweets it is considered viral).

More Read

Disease Management Programs – A Time to Revisit
Data Mining: Widespread Acceptance When?
Right Time Business Optimization
Design a Successful Business Dashboard: 3 Essential Tips
Why Human Resource Should Care About Big Data

The information that is given to the algorithm is the text of the Tweet and its author. Many other parameters can be taken into consideration such as the time that the Tweet has been posted, the type of the Tweet (e.g., politics, technology, health) or even whether this Tweet is part of a novel subject. Here is the output of the software that performs the predictions:


The number of Re-Tweets is shown in squares. Pay also close attention to the circled text shown above. For each Tweet the most probable outcome is given (‘t’= Tweet will become viral, ‘f’=otherwise) and a confidence for each prediction is given as a number from 0 to 1. As an example, the first Tweet shown above was posted from Paula Abdul saying that she will not return to American Idol. The algorithm predicts with a confidence of 63.38% that what Paula Abdul posted will be interesting (and it actually was).

The predictive model has an overall accuracy of 72.88% in predicting which Tweets will be viral in a total of 59 Tweets. An example of an incorrect prediction can be seen at the 4th circle from the top. The algorithm gave a 53.66% confidence that this Tweet will not become viral but actually this was a viral Tweet.

You can find the text file of the actual run from the algorithm here.

By looking the text file, results metrics such as TP (True positives) versus FP (False positives) can be calculated. It is also interesting to see how the algorithm switches to negative predictions when the number of Re-Tweets of each Tweet become less than 30.

Even though the example given here is very simplistic – and optimistic – the application of a tool of this kind for PR, marketing and branding could prove very useful. Marketeers can try different messages and see what impact each message is likely to have. Consider the following run that shows that @mashable is more influential than @lifeanalytics :


The following run shows that specific keywords raise our chances in making a Viral Tweet :

In theory this information could provide the basis for performing A/B tests: One could simply use the 2 messages shown above and record what impact each one has using Google Analytics (a process which could prove whether this technology works or not).

Finding information that is interesting to masses is actually a much harder problem. Twitter is a data source that is biased for many reasons: Specific people can pass their messages with great ease and Twitter is used by specific population segments. Almost a week ago I came across reddit and I believe that this site (and also Digg) is able to capture the preference of masses in a more efficient way than Twitter. The truth is that the available information from forums, blogs and many other websites can capture different aspects of human behavior. All that is needed to extract useful knowledge is an efficient blending of these facts, emotions and beliefs of people from different web sources.

Link to original post

TAGGED:text miningviral tweet
Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

image fx (2)
Monitoring Data Without Turning into Big Brother
Big Data Exclusive
image fx (71)
The Power of AI for Personalization in Email
Artificial Intelligence Exclusive Marketing
image fx (67)
Improving LinkedIn Ad Strategies with Data Analytics
Analytics Big Data Exclusive Software
big data and remote work
Data Helps Speech-Language Pathologists Deliver Better Results
Analytics Big Data Exclusive

Stay Connected

1.2kFollowersLike
33.7kFollowersFollow
222FollowersPin

You Might also Like

Twitter Analytics : These words may be affecting your popularity

6 Min Read

Please check out the Text Analytics Summit, Boston, June 12-13

2 Min Read

The Zero Latency Future is Now

5 Min Read
Text Mining
Big DataData Mining

Text Mining Strategies and Limitations with Scalable Data Solutions

5 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

AI and chatbots
Chatbots and SEO: How Can Chatbots Improve Your SEO Ranking?
Artificial Intelligence Chatbots Exclusive
giveaway chatbots
How To Get An Award Winning Giveaway Bot
Big Data Chatbots Exclusive

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-25 SmartData Collective. All Rights Reserved.
Go to mobile version
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?