Hands-On Big Data Analytics with PySpark: Analyze large datasets and discover techniques for testing, immunizing, and parallelizing Spark jobs
Book information
Description
Use PySpark to easily crush messy data at-scale and discover proven techniques to create testable, immutable, and easily parallelizable Spark jobs Key FeaturesWork with large amounts of agile data using distributed datasets and in-memory cachingSource data from all popular data hosting platforms, such as HDFS, Hive, JSON, and S3Employ the easy-to-use PySpark API to deploy big data Analytics for production Book Description Apache Spark is an open source parallel-processing framework that has been around for quite some time now. One of the many uses of Apache Spark is for data analytics applications across clustered computers. In this book, you will not only learn how to use Spark and the Python API to create high-performance analytics with big data, but also discover techniques for testing, immunizing, and parallelizing Spark jobs. You will learn how to source data from all popular data hosting platforms, including HDFS, Hive, JSON, and...
Similar books
Il Papa Non Eletto: Giuseppe Siri cardinale di Santa Romana Chiesa
1993 · PDF
Il Papa Non Eletto: Giuseppe Siri cardinale di Santa Romana Chiesa
1993 · PDF
Il Papa Non Eletto: Giuseppe Siri cardinale di Santa Romana Chiesa
1993 · EPUB
Die Paradoxie des Rechts
2014 · PDF
Hong Kong 1941–45: First strike in the Pacific War
White Tigress Green Dragon: Taoist Sexual Secrets for Youthful Restoration and Spiritual Illumination
2015 · AZW3
The Sexual Teachings of the Jade Dragon: Taoist Methods for Male Sexual Revitalization
EPUB
Hong Kong, 1941-45: first strike in the Pacific War
2014 · EPUB