ENGLISH

Beginning Apache Spark Using Azure Databricks: Unleashing Large Cluster Analytics in the Cloud

Book information

Publisher
Apress
Year
2020
ISBN
1484257804, 9781484257807
Language
english
Format
PDF
Filesize
3 MB (2939212 bytes)
Edition
1
Pages
300\281
Time added
2020-06-11 17:18:32

Description

Analyze vast amounts of data in record time using Apache Spark with Databricks in the Cloud. Learn the fundamentals, and more, of running analytics on large clusters in Azure and AWS, using Apache Spark with Databricks on top. Discover how to squeeze the most value out of your data at a mere fraction of what classical analytics solutions cost, while at the same time getting the results you need, incrementally faster. This book explains how the confluence of these pivotal technologies gives you enormous power, and cheaply, when it comes to huge datasets. You will begin by learning how cloud infrastructure makes it possible to scale your code to large amounts of processing units, without having to pay for the machinery in advance. From there you will learn how Apache Spark, an open source framework, can enable all those CPUs for data analytics use. Finally, you will see how services such as Databricks provide the power of Apache Spark, without you having to know anything about configuring hardware or software. By removing the need for expensive experts and hardware, your resources can instead be allocated to actually finding business value in the data. This book guides you through some advanced topics such as analytics in the cloud, data lakes, data ingestion, architecture, machine learning, and tools, including Apache Spark, Apache Hadoop, Apache Hive, Python, and SQL. Valuable exercises help reinforce what you have learned. What You Will Learn Discover the value of big data analytics that leverage the power of the cloudGet started with Databricks using SQL and Python in either Microsoft Azure or AWSUnderstand the underlying technology, and how the cloud and Apache Spark fit into the bigger picture See how these tools are used in the real world Run basic analytics, including machine learning, on billions of rows at a fraction of a cost or free Who This Book Is For Data engineers, data scientists, and cloud architects who want or need to run advanced analytics in the cloud. It is assumed that the reader has data experience, but perhaps minimal exposure to Apache Spark and Azure Databricks. The book is also recommended for people who want to get started in the analytics field, as it provides a strong foundation. Table of Contents About the Author About the Technical Reviewer Introduction Chapter 1: Introduction to Large-Scale Data Analytics Analytics, the hype Analytics, the reality Large-scale analytics for fun and profit Data: Fueling analytics Free as in speech. And beer! Into the clouds Databricks: Analytics for the lazy ones How to analyze data Large-scale examples from the real world Telematics at Volvo Trucks Fraud detection at Visa Customer analytics at Target Targeted ads at Cambridge Analytica Summary Chapter 2: Spark and Databricks Apache Spark, the short overview Databricks: Managed Spark The far side of the Databricks moon Spark architecture Apache Spark processing Working with data Data processing Storing data Cool components on top of Core Summary Chapter 3: Getting Started with Databricks Cloud-only Community edition: No money? No problem Mostly good enough Getting started with the community edition Commercial editions: The ones you want Databricks on Amazon Web Services Azure Databricks Summary Chapter 4: Workspaces, Clusters, and Notebooks Getting around in the UI Clusters: Powering up the engines Data: Getting access to the fuel Notebooks: Where the work happens Summary Chapter 5: Getting Data into Databricks Databricks File System Navigating the file system The FileStore, a portal to your data Schemas, databases, and tables Hive Metastore The many types of source files Going binary Alternative transportation Importing from your computer Getting data from the Web Working with the shell Basic importing with Python Getting data with SQL Mounting a file system Mounting example Amazon S3 Mounting example Microsoft Blog Storage Getting rid of the mounts How to get data out of Databricks Summary Chapter 6: Querying Data Using SQL The Databricks flavor Getting started Picking up data Filtering data Joins and merges Ordering data Functions Windowing functions A view worth keeping Hierarchical data Creating data Manipulating data Delta Lake SQL UPDATE, DELETE, and MERGE Keeping Delta Lake in order Transaction logs Selecting metadata Gathering statistics Summary Chapter 7: The Power of Python Python: The language of choice A turbo-charged intro to Python Finding the data DataFrames: Where active data lives Getting some data Selecting data from DataFrames Chaining combo commands Working with multiple DataFrames Slamming data together Summary Chapter 8: ETL and Advanced Data Wrangling ETL: A recap An overview of the Spark UI Cleaning and transforming data Finding nulls Getting rid of nulls Filling nulls with values Removing duplicates Identifying and clearing out extreme values Taking care of columns Pivoting Explode When being lazy is good Caching data Data compression A short note about functions Lambda functions Storing and shuffling data Save modes Managed vs. unmanaged tables Handling partitions Summary Chapter 9: Connecting to and from Databricks Connecting to and from Databricks Getting ODBC and JDBC up and running Creating a token Preparing the cluster Let’s create a test table Setting up ODBC on Windows Setting up ODBC on OS X Connecting tools to Databricks Microsoft Excel on Windows Microsoft Power BI Desktop on Windows Tableau on OS X PyCharm (and more) via Databricks Connect Using RStudio Server Accessing external systems A quick recap of libraries Connecting to external systems Azure SQL Oracle MongoDB Summary Chapter 10: Running in Production General advice Assume the worst Write rerunnable code Document in the code Write clear, simple code Print relevant stuff Jobs Scheduling Running notebooks from notebooks Widgets Running jobs with parameters The command line interface Setting up the CLI Running CLI commands Creating and running jobs Accessing the Databricks File System Picking up notebooks Keeping secrets Secrets with privileges Revisiting cost Users, groups, and security options Users and groups Using SCIM provisioning Access Control Workspace Access Control Cluster, Pool, and Jobs Access Control Table Access Control Personal Access Tokens The rest Summary Chapter 11: Bits and Pieces MLlib Frequent Pattern Growth Creating some data Preparing the data Running the algorithm Parsing the results MLflow Running the code Checking the results Updating tables Create the original table Connect from Databricks Pulling the delta Verifying the formats Update the table A short note about Pandas Koalas, Pandas for Spark Playing around with Koalas The future of Koalas The art of presenting data Preparing data Using Matplotlib Building and showing the dashboard Adding a widget Adding a graph Schedule run REST API and Databricks What you can do What you can’t do Getting ready for APIs Example: Get cluster data Example: Set up and execute a job Example: Get the notebooks All the APIs and what they do Delta streaming Running a stream Checking and stopping the streams Running it faster Using checkpoints Index

Similar books

Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36

Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36

2010 · PDF

THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.

THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.

1858 · PDF

Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.

Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.

2022 · PDF