{"product_id":"154699","title":"Running Spark ","description":"\u003ccenter\u003e\u003cdiv style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/tmgdisk01.cafe24.com\/images\/vs\/4172\/sv\/3jYDPGxXxrPeoAh75YnA3KOWD2l4KU.png?v=1765081420\" style=\"max-width:100%;max-height:10px\"\u003e\u003c\/div\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\n\n\u003cdiv style=\"width:95%\"\u003e\n\n\u003cdiv style=\"text-align:center;font-size:30px;font-weight:bolder;line-height:1.6em\"\u003e Running Spark \u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"border-bottom:1px;border-bottom-style:dotted;border-color:;padding-bottom:20px\"\u003e\u003ccenter\u003e\u003ctable align=\"center\" width=\"100%\"\u003e\u003ctbody style=\"border:0px\"\u003e\n\n\u003ctr\u003e\u003ctd align=\"center\" style=\"line-height:1.2em;text-align:center;font-size:18px;color:black;font-weight:bold;padding-bottom:20px;\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\u003ctr\u003e\u003ctd style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/image.yes24.com\/goods\/110088102\/XL\" style=\"max-width:100%;height:auto\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\n\u003c\/tbody\u003e\u003c\/table\u003e\u003c\/center\u003e\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;{split_style6}padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e Description \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;word-break:break-all;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eBook Introduction\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\u003cdiv\u003e \u003cb\u003eThe definitive Spark introductory book recommended by Spark founder Matei Zaharia!\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e This updated edition, which includes Spark 3.x, shows data engineers and data scientists why Spark's architecture and integration are important.\u003cbr\u003e We perform data analysis from simple to complex and systematically explain how to use machine learning algorithms.\u003cbr\u003e\n\n\u003c\/div\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cul\u003e\u003cli\u003e You can preview some of the book's contents.\u003cbr\u003e \u003cspan\u003ePreview\u003c\/span\u003e\n\n\u003c\/li\u003e\u003c\/ul\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eindex\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e Translator's Preface x\u003cbr\u003e Beta Reader Review xii\u003cbr\u003e Recommendation xiv\u003cbr\u003e Starting with xv\u003cbr\u003e About the cover xxi\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 1: Introducing Apache Spark: A Unified Analytics Engine\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e The Beginning of Spark 1\u003cbr\u003e What is Apache Spark? 4\u003cbr\u003e Integrated Analysis 7 \u003cbr\u003eDeveloper Experience 15\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 2 Downloading and Starting Apache Spark 19\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Step 1: Download Apache Spark 19\u003cbr\u003e Step 2: Using the Scala or PySpark Shell 22\u003cbr\u003e Using a local machine 24\u003cbr\u003e Step 3: Understanding Spark Application Concepts 26\u003cbr\u003e Transformation, Action, and Delayed Evaluation 29\u003cbr\u003e Spark UI 31\u003cbr\u003e First standalone application 34\u003cbr\u003e Summary 42\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 3 Apache Spark's Formalized API 43\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Spark: What's Underneath an RDD? 44\u003cbr\u003e Establishing Spark's Structure 45\u003cbr\u003e DataFrame API 48\u003cbr\u003e Dataset API 71\u003cbr\u003e DataFrames vs. DataSets 77\u003cbr\u003e Spark SQL and the underlying engine 79\u003cbr\u003e Summary 85\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 4 Spark SQL and DataFrames: Introducing Built-in Data Sources 86\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Using Spark SQL in Spark Applications 87\u003cbr\u003e SQL Tables and Views 93\u003cbr\u003e Data Sources for DataFrames and SQL Tables 98\u003cbr\u003e Summary 119\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 5 Spark SQL and DataFrames: Interacting with External Data Sources 120\u003c\/b\u003e\u003cbr\u003e \u003cbr\u003eSpark SQL and Apache Hive 120\u003cbr\u003e Querying with Spark SQL Shell, Beeline, and Tableau 126\u003cbr\u003e External Data Source 134\u003cbr\u003e PostgreSQL 137\u003cbr\u003e Higher-Order Functions in DataFrames and Spark SQL 144\u003cbr\u003e 150 Common DataFrame and Spark SQL Operations\u003cbr\u003e Summary 163\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 6 Spark SQL and Datasets 164\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e A Single API for Java and Scala 164\u003cbr\u003e Working with Datasets 167\u003cbr\u003e Memory Management for Datasets and DataFrames 175\u003cbr\u003e Dataset Encoder 176\u003cbr\u003e Dataset usage cost 178\u003cbr\u003e Summary 180\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 7 Optimizing and Tuning Spark Applications 181\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Optimizing and Tuning Spark Efficiently 181\u003cbr\u003e Data Caching and Persistence 191\u003cbr\u003e Spark Join Types 196\u003cbr\u003e A Look Inside Spark UI 206\u003cbr\u003e Summary 213\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 8: Streaming Standardization 214\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Evolution of Apache Spark's Stream Processing Engine 214\u003cbr\u003e Programming Model for Formatted Streaming 218\u003cbr\u003e Fundamentals of Structured Streaming Queries 220\u003cbr\u003e Internals of a Running Streaming Query 227 \u003cbr\u003eStreaming Data Sources and Sinks 233\u003cbr\u003e Data Transformation 243\u003cbr\u003e Streaming Aggregation with Stateful Information 246\u003cbr\u003e Streaming Join 255\u003cbr\u003e Arbitrary state maintenance operations 263\u003cbr\u003e Performance Tuning 272\u003cbr\u003e Summary 274\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 9 Building a Reliable Data Lake with Apache Spark 275\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e The Importance of Optimal Storage Solutions 275\u003cbr\u003e Database 277\u003cbr\u003e Data Lake 279\u003cbr\u003e Lakehouse: The Next Step in the Evolution of Storage Solutions 282\u003cbr\u003e Building a Lakehouse with Apache Spark and Delta Lake 285\u003cbr\u003e Summary 296\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 10 Machine Learning with MLlib 298\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e What is Machine Learning? 299\u003cbr\u003e Machine Learning Pipeline Design 302\u003cbr\u003e Hyperparameter Tuning 322\u003cbr\u003e Summary 338\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 11: Managing, Deploying, and Scaling Machine Learning Pipelines with Apache Spark 339\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Model Management 339\u003cbr\u003e Model Deployment Options with MLlib 346\u003cbr\u003e Leveraging Spark for Non-MLlib Models 352\u003cbr\u003e Summary 358\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 12 Epilogue: Apache Spark 3.0 359\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Spark Core and Spark SQL 359 \u003cbr\u003eJeonghyeonghwa Streaming 368\u003cbr\u003e PySpark, Pandas UDF, Pandas Function API 370\u003cbr\u003e Changed features 373\u003cbr\u003e Summary 376\u003cbr\u003e\u003cbr\u003e Search 379\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eDetailed image\u003c\/b\u003e \u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cdiv\u003e\u003cimg src=\"https:\/\/image.yes24.com\/momo\/TopCate3859\/MidCate009\/385888852.jpg\" border=\"0\" alt=\"Detailed Image 1\"\u003e\u003c\/div\u003e\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eInto the book\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e By 2013, Spark had become so widely used that its original authors and researchers (Matej Zaharia, Ali Ghosh, Reynolds Shin, Patrick Wendell, Ion Stojka, and Andy Konvinski) transferred the Spark project to the Apache Software Foundation (ASF) and formed a company called Databricks.\u003cbr\u003e Developers from Databricks and the open source community released Apache Spark 1.010 in May 2014 under the leadership of the ASF.\u003cbr\u003e This first major release sets the stage for frequent releases and notable features from Databricks and over 100 commercial partners.\u003cbr\u003e\u003cbr\u003e --- p.4\u003cbr\u003e \u003cbr\u003eYou can write a single Spark application and everything runs on it, without having to spin up separate engines for completely different tasks or learn separate APIs.\u003cbr\u003e With Spark, you have a single, unified processing engine to handle your workload.\u003cbr\u003e\u003cbr\u003e --- p.5\u003cbr\u003e\u003cbr\u003e Of all the joys a developer can experience, one of the most appealing is a well-structured set of APIs that increase productivity, are easy to use, and are easy to understand.\u003cbr\u003e One of the principles of Apache Spark is to appeal to developers with an easy-to-use API across multiple languages, including Scala, Java, Python, SQL, and R, regardless of the scale of the data.\u003cbr\u003e\u003cbr\u003e --- p.15\u003cbr\u003e \u003cbr\u003eOne of the authors of this book is a data scientist who loves baking cookies using M\u0026amp;Ms, and she often gives them out as prizes to students from various states in her machine learning and data science courses.\u003cbr\u003e But because she's a data-driven person, she wants to make sure that students in different states are given the right color of M\u0026amp;Ms.\u003cbr\u003e Let's write a Spark program that reads a file containing over 100,000 data points (each line contains a state, a color, and a count of M\u0026amp;Ms) and aggregates them by color and state.\u003cbr\u003e These aggregated results will tell you what color M\u0026amp;Ms students in each state like.\u003cbr\u003e The complete Python program is in Example 2-1.\u003cbr\u003e\u003cbr\u003e --- p.35\u003cbr\u003e\u003cbr\u003e What's the difference between data caching and persistence? In Spark, the two terms are synonymous. \u003cbr\u003eTwo API calls, cache() and persist(), provide these features.\u003cbr\u003e The latter can provide more granular settings about where and how data is stored - whether in memory or on disk, and whether it is serialized or not.\u003cbr\u003e\n\n\u003c\/div\u003e\n\u003cdiv\u003e --- p.191\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003ePublisher's Review\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e \u003cb\u003eThe definitive Spark introductory book recommended by Spark founder Matei Zaharia!\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e The second edition of Running Spark: Lightning-Fast Data Analysis has been published.\u003cbr\u003e As data grows larger, is generated faster, and is available in a variety of formats, large-scale processing for analytics and machine learning is also required.\u003cbr\u003e Apache Spark is an alternative that can efficiently handle these large-scale workloads.\u003cbr\u003e\u003cbr\u003e This updated edition, which includes Spark 3.x, shows data engineers and data scientists why Spark's architecture and integration are important. \u003cbr\u003eWe systematically explain how to perform data analysis from simple to complex and how to use machine learning algorithms.\u003cbr\u003e Step-by-step exercises, code examples, notebooks, and more help you:\u003cbr\u003e\u003cbr\u003e ■ Learning high-level structured APIs using Python, SQL, Scala, and Java\u003cbr\u003e ■ Understanding Spark Jobs and the SQL Engine\u003cbr\u003e ■ Inspect, tune, and debug Spark jobs using Spark configuration and Spark UI\u003cbr\u003e ■ Connect to data sources such as JSON, Parque, CSV, Avro, ORC, Hive, S3, and Kafka\u003cbr\u003e ■ Perform analytics on batch and streaming data using structured streaming\u003cbr\u003e ■ Building a stable data pipeline with open source Delta Lake and Spark\u003cbr\u003e ■ Develop machine learning pipelines using MLlib and reproduce and deploy models using MLflow. \u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e GOODS SPECIFICS \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003ePublication date:\u003c\/strong\u003e June 24, 2022\u003c\/div\u003e\n\n \u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e- \u003cstrong\u003ePage count, weight, size:\u003c\/strong\u003e 404 pages | 788g | 188*257*20mm\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN13:\u003c\/strong\u003e 9791191600889\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN10:\u003c\/strong\u003e 1191600882 \u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cspan\u003e\u003c\/span\u003e\n\n\u003c\/center\u003e\n\n\n\u003c\/center\u003e","brand":"LIBRAIRIE COREENNE","offers":[{"title":"Default Title","offer_id":43893444116522,"sku":"154699","price":40.0,"currency_code":"EUR","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0683\/2750\/5962\/files\/c10571a351ab343802beaf575e274bb0.jpg?v=1765402388","url":"https:\/\/librairie.coreenne.fr\/en\/products\/154699","provider":"LIBRAIRIE COREENNE","version":"1.0","type":"link"}