We use cookies, including third-party cookies from Google to serve personalized ads through AdSense, to operate this site and understand how it is used. By continuing to browse, you accept this use. See our Privacy Policy and Terms of Use for details, including how to opt out of personalized advertising.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    chatgpt image jul 21, 2026, 04 34 30 pm
    4 Core Benefits of Predictive Maintenance after Vibration Analysis
    10 Min Read
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results -- AI-generated illustration
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results
    11 Min Read
    chatgpt image jul 13, 2026, 04 23 45 pm
    How Data Analytics Helps Companies Improve User Engagement
    19 Min Read
    chatgpt image jul 13, 2026, 03 59 46 pm
    How Data Analytics Improves Multi-Location Search Strategies
    10 Min Read
    cybersecurity efforts
    How Behavioral Analytics and AI Are Redefining Cybersecurity for Boca Raton Businesses
    14 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: Training students on mega-scale data
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Uncategorized > Training students on mega-scale data
Uncategorized

Training students on mega-scale data

DavidMSmith
DavidMSmith
3 Min Read
Training students on mega-scale data
Illustration generated with FLUX.2 [klein 4B] via Cloudflare Workers AI.
SHARE

In a New York Times article (sub. req.) published on the weekend, IBM and Google expressed doubts that the students graduating from US universities today have the chops to deal with the mulit-terabyte datasets that are becoming commonplace online and in domains like bioscience and astronomy today. From the article:

For the most part, university students have used rather modest computing systems to support their studies. They are learning to collect and manipulate information on personal computers or what are known as clusters, where computer servers are cabled together to form a larger computer. But even these machines fail to churn through enough data to really challenge and train a young mind meant to ponder the mega-scale problems of tomorrow.

The article reveals how Google and IBM are promoting internet-scale research at places like the University of Washington and Purdue. But a curious omission from the article is any mention of open-source technologies that are spurring the innovation in processing and analyzing these data sets. Tools like Hadoop, for processing internet-scale data sets and R, for analyzing the processed data (most likely in some parallelized form), and other open-source projects not yet conceived, are going to be critical in this endeavour.

New York Times: Training to Climb an Everest of Digital Data


More Read

Largest HIPAA Breach Ever: Hackers Steal Data on 4.5 Million Community Health Systems Patients
Largest HIPAA Breach Ever: Hackers Steal Data on 4.5 Million Community Health Systems Patients
SnapLogic: Making Big Data Integration as a Service a Hadoop Reality
Big Data Sets You Can Use with R
Which Big Data Personality Are You? [INFOGRAPHIC]
Cloud ROI: Business and Financial Gains
TAGGED:googlehadoopibmnew york timesr
Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

Synthetic Data vs Real Web Data: Comparison, Limitations, and Collection Methods  -- AI-generated illustration
Synthetic Data vs Real Web Data: Comparison, Limitations, and Collection Methods 
Big Data Exclusive
Illustration of mobile analytics dashboards with ad performance charts connected to backend databases
11 Best Sisense Alternatives for Embedded Analytics
Business Intelligence Exclusive
Analyst points at colorful circular data dashboard on screen - information technology business metrics
How Fragmented Workplace Tech Undermines Reliable Business Metrics and Reporting
Cloud Computing Exclusive Infographic IT
Using Multi-Source Data and Analytics to Detect Operational Drift Across Franchise Networks -- AI-generated illustration
Using Multi-Source Data and Analytics to Detect Operational Drift Across Franchise Networks
Exclusive Infographic

Stay Connected

1.2KFollowersLike
33.7KFollowersFollow
222FollowersPin

You Might also Like

Experimenting on Facebook
Data MiningPredictive Analytics

Experimenting on Facebook

5 Min Read
The Fallacy of the Data Scientist Shortage
AnalyticsBusiness IntelligenceBusiness RulesCloud ComputingCollaborative DataCommentaryData MiningData WarehousingDecision ManagementHadoopJobsMapReducePredictive AnalyticsR Programming LanguageSentiment AnalyticsStatisticsText AnalyticsUnstructured Data

The Fallacy of the Data Scientist Shortage

8 Min Read
Google+ Is After Your Friends with Big Data and Beautiful Photos
AnalyticsBig DataExclusiveSocial DataSocial Media Analytics

Google+ Is After Your Friends with Big Data and Beautiful Photos

5 Min Read
Memo to Steve Ballmer: Just Ask Them!
Data Mining

Memo to Steve Ballmer: Just Ask Them!

4 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

From Bolts to Bots: How AI Is Fortifying the Automotive Industry
From Bolts to Bots: How AI Is Fortifying the Automotive Industry
Artificial Intelligence
How To Get An Award Winning Giveaway Bot
How To Get An Award Winning Giveaway Bot
Big Data Chatbots Exclusive

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-26 SmartData Collective. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?