Fast Data Processing with Spark
Book information
Description
Spark is a framework for writing fast, distributed programs. Spark solves similar problems as Hadoop MapReduce does but with a fast in-memory approach and a clean functional style API. With its ability to integrate with Hadoop and inbuilt tools for interactive query analysis (Shark), large-scale graph processing and analysis (Bagel), and real-time analysis (Spark Streaming), it can be interactively used to quickly process and query big data sets. Fast Data Processing with Spark covers how to write distributed map reduce style programs with Spark. The book will guide you through every step required to write effective distributed programs from setting up your cluster and interactively exploring the API, to deploying your job to the cluster, and tuning it for your purposes.
Similar books
High Performance Spark: Best Practices for Scaling and Optimizing Apache Spark
2017 · EPUB
Scaling Python with Dask: From Data Science to Machine Learning
2023 · PDF
Scaling Python with Dask: From Data Science to Machine Learning
2023 · EPUB
Scaling Python with Dask: From Data Science to Machine Learning
2023 · EPUB
Scaling Python with Ray: Adventures in Cloud and Serverless Patterns
EPUB
Scaling Python with Dask: From Data Science to Machine Learning
MOBI
Scaling Python with Ray: Adventures in Cloud and Serverless Patterns
2023 · PDF
Kubeflow for Machine Learning: From Lab to Production
2020 · EPUB