Showing posts with label Hadoop. Show all posts
Showing posts with label Hadoop. Show all posts

Wednesday, 3 February 2016

Difference between Hadoop and Apache Spark

Hadoop and Apache Spark are seen as the competitors in the world of big data, but now the growing consensus is that they are better convention in together. Here is a brief look at what they do and how they are compared.  1. They do different things: Both are the big-data frameworks, but they do not serve the same purposes. Hadoop is a distributed data infrastructure. It also Indexes and keep track of that data, enabling big-data processing and analytics. On the other hand, Spark is a data processing tool. Secondly, both can be used individually, without the other. 3. Spark is faster 4. You may not need Spark's speed: Spark is fit for real-time marketing campaigns, online product recommendations, cybersecurity analytics and machine log monitoring. 5. Failure recovery: differently, but still good. Read more at: http://www.computerworld.com/article/3014516/big-data/5-things-to-know-about-hadoop-v-apache-spark.html

Wednesday, 6 January 2016

The most common data science skills

As the field of Data Science is growing, the confusion regarding the skills needed to be a data scientist is also increasing. Most of us think data science skills range from computer science and statistics, to machine learning and strong communication. But, the top data science skills list includes data analysis at the top, followed by others like R, Python and machine learning. As per recruiter lists, R, Python, SQL, SAS and Hadoop are appreciated. To know more about data science skills, follow the article written by Daniel Levine (Content Marketer for RJMetrics) at: http://www.smartdatacollective.com/daniellevine/366486/top-20-data-science-skills