{"product_id":"139507","title":"Technology that supports big data ","description":"\u003ccenter\u003e\u003cdiv style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/tmgdisk01.cafe24.com\/images\/vs\/4172\/sv\/3jXPBvfsZhfROTM72HQV3L7S7pu7Tr.png?v=1765075179\" style=\"max-width:100%;max-height:10px\"\u003e\u003c\/div\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\n\n\u003cdiv style=\"width:95%\"\u003e\n\n\u003cdiv style=\"text-align:center;font-size:30px;font-weight:bolder;line-height:1.6em\"\u003e Technology that supports big data \u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"border-bottom:1px;border-bottom-style:dotted;border-color:;padding-bottom:20px\"\u003e\u003ccenter\u003e\u003ctable align=\"center\" width=\"100%\"\u003e\u003ctbody style=\"border:0px\"\u003e\n\n\u003ctr\u003e\u003ctd align=\"center\" style=\"line-height:1.2em;text-align:center;font-size:18px;color:black;font-weight:bold;padding-bottom:20px;\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\u003ctr\u003e\u003ctd style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/image.yes24.com\/goods\/66277191\/XL\" style=\"max-width:100%;height:auto\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\n\u003c\/tbody\u003e\u003c\/table\u003e\u003c\/center\u003e\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;{split_style6}padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e Description \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;word-break:break-all;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eBook Introduction\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\u003cdiv\u003e \u003cb\u003eThe success of modern business hinges on how data is collected, integrated, and processed!\u003cbr\u003e Everything you need to know about big data and related technologies, from a data processing expert!\u003cbr\u003e 'How to systematize data processing?'\u003c\/b\u003e\u003cbr\u003e \u003cbr\u003e\"Technology Supporting Big Data\" organizes the elements and technologies required for a series of data processing, focusing on engineering issues, establishes a foundation for efficient data processing, and covers various technologies that support system automation on top of that.\u003cbr\u003e As computer performance improves, expectations for the development of data-driven systems, led by machine learning, are growing.\u003cbr\u003e Therefore, in the future, regardless of the system size, the demand for 'technology that makes data processing itself part of the system' will gradually increase.\u003cbr\u003e The diverse visual materials and systematic introduction to related technologies presented in this book will be of great help to readers in their introduction to big data.\u003cbr\u003e\n\n\u003c\/div\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cul\u003e\u003cli\u003e You can preview some of the book's contents.\u003cbr\u003e \u003cspan\u003ePreview\u003c\/span\u003e\n\n\u003c\/li\u003e\u003c\/ul\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eindex\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e CHAPTER 1 Basic Knowledge of Big Data _ 1\u003cbr\u003e 1-1 [Background] Big Data Establishment 3 \u003cbr\u003eAccelerated data processing through distributed systems: Two representative technologies that overcome the difficulty of handling big data.\u003cbr\u003e Pioneering Business Use of Distributed Systems - Coexisting with Data Warehouses 7\u003cbr\u003e Expanding the Breadth of Data Analysis You Can Do Yourself - Accelerate Big Data Utilization with Cloud Services and Data Discovery 8\u003cbr\u003e 1-2 Data Analysis Foundation in the Big Data Era 11\u003cbr\u003e [Re-introduction] Big Data Technology - Data Processing Structure Using Distributed Systems 11\u003cbr\u003e Data Warehouses and Data Marts - Data Pipeline Basics 16\u003cbr\u003e Data Lakes - Accumulating Data as It Is 17\u003cbr\u003e Developing a Data Analytics Foundation Step by Step - Teams and Roles, Starting Small and Expanding 19\u003cbr\u003e Purpose of Data Collection - Three Examples: 'Retrieval,' 'Processing,' and 'Visualization' 22\u003cbr\u003e Confirmatory and Exploratory Data Analysis 25 \u003cbr\u003e1-3 [Attribute Learning] Special Analysis and Data Frames Using Scripting Languages ​​26\u003cbr\u003e Data Processing and Scripting Languages ​​- Popular Python Languages ​​and DataFrames 26\u003cbr\u003e Data Frames, the Basics of the Basics - Building from 'Arrays within Arrays' 27\u003cbr\u003e Example of a web server's access log - easily processed with pandas data frames 28\u003cbr\u003e Interactively Aggregating Time Series Data - Aggregating Data Using DataFrames 30\u003cbr\u003e Using SQL Results as a Data Frame 31\u003cbr\u003e 1-4 BI Tools and Monitoring 33\u003cbr\u003e Monitoring with Spreadsheets - Understanding the Current Status of Your Project 33\u003cbr\u003e Data-Driven Decision Making - KPI Monitoring 35\u003cbr\u003e Identifying Changes and Understanding the Details - Using BI Tools 37\u003cbr\u003e Determining the Line Between Manual and Automated Work 39\u003cbr\u003e 1-5 Summary 42\u003cbr\u003e\u003cbr\u003e CHAPTER 2 Exploring Big Data _ 43\u003cbr\u003e Basic 45 of 2-1 cross tally \u003cbr\u003eTransaction Tables, Cross Tables, and Pivot Tables—The Concept of 'Cross Aggregation' 45\u003cbr\u003e Lookup Tables - Adding Attributes by Combining Tables 47\u003cbr\u003e Aggregating Tables with SQL - Preparing for Cross-Aggregation of Bulk Data 50\u003cbr\u003e Data Aggregation? Data Mart? Visualization—System Configuration Determined by the Size of the Data Mart 55\u003cbr\u003e 2-2 High-speed through column-oriented storage 56\u003cbr\u003e Reducing database latency 56\u003cbr\u003e Column-Oriented Database Access - Compressing Columns to Reduce Disk I\/O 58\u003cbr\u003e An MPP Database Approach - Leveraging Multi-Core Power through Parallelism 61\u003cbr\u003e 2-3 Ad hoc analysis and visualization tools 64\u003cbr\u003e Ad-hoc Analysis with Jupyter Notebooks - Recording Analysis Processes in a Notebook 64\u003cbr\u003e Dashboard Tools - Visualize aggregated results regularly 68\u003cbr\u003e BI Tools - Interactive Dashboards 75\u003cbr\u003e 2-4 Basic Structure of a Data Mart 77 \u003cbr\u003eBuilding Data Marts Perfect for Visualization - OLAP 77\u003cbr\u003e Denormalizing a Table 79\u003cbr\u003e Abstracting Tables in Preparation for Multidimensional Model Visualization 82\u003cbr\u003e 2-5 Summary 86\u003cbr\u003e\u003cbr\u003e CHAPTER 3 Distributed Processing of Big Data _ 87\u003cbr\u003e 3-1 Framework for Large-Scale Distributed Processing 89\u003cbr\u003e Structured and Unstructured Data 89\u003cbr\u003e Hadoop - A Common Platform for Distributed Data Processing 92\u003cbr\u003e Spark - High-Speed ​​In-Memory Data Processing 99\u003cbr\u003e 3-2 Query Engine 101\u003cbr\u003e Pipeline 101: Building a Data Mart\u003cbr\u003e Creating Structured Data with Hive 102\u003cbr\u003e The Structure of Presto, the Interactive Query Engine - Aggregating Structured Data with Presto 109\u003cbr\u003e Choosing a Data Analysis Framework: MPP Databases, Hive, Presto, and Spark 115\u003cbr\u003e 3-3 Building a Data Mart 119\u003cbr\u003e Fact Tables - Accumulating Time Series Data 119\u003cbr\u003e Aggregate Tables - Reducing the Number of Records 122\u003cbr\u003e Snapshot Table - Recording the Master's State 123\u003cbr\u003e History Table - Recording Master Changes 127 \u003cbr\u003e[Final Step] Complete the Denormalized Table by Adding Dimensions 127\u003cbr\u003e 3-4 Summary 130\u003cbr\u003e\u003cbr\u003e CHAPTER 4 Accumulation of Big Data _ 131\u003cbr\u003e 4-1 Bulk and Streaming Data Collection 133\u003cbr\u003e Object Storage and Data Ingestion - Loading Data into Distributed Storage 133\u003cbr\u003e Bulk Data Transfer - The Need for an ETL Server 135\u003cbr\u003e Streaming data transfer - Data transfer for handling small data that is continuously transmitted 137\u003cbr\u003e 4-2 [Performance × Reliability] Tradeoffs in Message Delivery 143\u003cbr\u003e Message Broker - Installing a Middle Layer to Solve Storage Performance Problems 143\u003cbr\u003e Ensuring Message Delivery is Difficult - Reliability Issues and Three Design Approaches 146\u003cbr\u003e Deduplication is a costly operation 149\u003cbr\u003e Data Collection Pipelines - Storage Suitable for Long-Term Data Analysis 152\u003cbr\u003e 4-3 Optimization of Time Series Data 154 \u003cbr\u003eProcess time and event time - The main target of data analysis is event time 154\u003cbr\u003e Partitioning and Problems by Process Time - Full Scans You Want to Avoid as Much as Possible 154\u003cbr\u003e Time Series Indexes - Efficient Aggregation by Event Time ① 156\u003cbr\u003e Conditional Pushdown - Efficient Aggregation by Event Time ② 157\u003cbr\u003e Partitioning by Event Time - Table Partitioning, Time Series Tables 158\u003cbr\u003e 4-4 Distributed Storage of Unstructured Data 161\u003cbr\u003e [Basic Strategy] Data Utilization with NoSQL Databases 161\u003cbr\u003e Distributed KVS - Improving Disk Write Performance 162\u003cbr\u003e Wide Column Store - Analyzing and Storing Structured Data 166\u003cbr\u003e Document Store - Managing Schemaless Data 169\u003cbr\u003e Search Engines - Finding Data with Keyword Searches 171\u003cbr\u003e 4-5 Summary 175\u003cbr\u003e\u003cbr\u003e CHAPTER 5 Big Data Pipeline _ 177\u003cbr\u003e 5-1 Workflow Management 179\u003cbr\u003e [Basic Knowledge] Workflow Management - Managing the Flow of Data 179 \u003cbr\u003eThinking First About How to Recover from Errors 183\u003cbr\u003e Describing tasks as idempotent operations - executing the same task multiple times produces the same result 188\u003cbr\u003e Making the entire workflow idempotent 194\u003cbr\u003e Task Queues - Controlling Resource Consumption 195\u003cbr\u003e 5-2 Batch-type data flow 199\u003cbr\u003e The Era of MapReduce Is Over - Data Flow and Workflow 199\u003cbr\u003e A New Framework to Replace MapReduce - Internal Representation via DAGs 201\u003cbr\u003e Combining Data Flows and Workflows 204\u003cbr\u003e Separating Data Flow and SQL - Data Warehouse Pipelines and Data Mart Pipelines 207\u003cbr\u003e 5-3 Streaming Data Flow 209\u003cbr\u003e Splitting the Path between Batch and Stream Processing 209\u003cbr\u003e Integrating Batch and Stream Processing 211\u003cbr\u003e Replacing the Results of Stream Processing with Batch Processing - Addressing Two Problems of Stream Processing 214\u003cbr\u003e Out-of-order data processing 217 \u003cbr\u003e5-4 Summary 220\u003cbr\u003e\u003cbr\u003e CHAPTER 6 Building a Big Data Analysis Foundation _ 223\u003cbr\u003e 6-1 Ad Hoc Analysis of Schemaless Data 225\u003cbr\u003e Collecting Schemaless Data 225\u003cbr\u003e Preparing the Interactive Execution Environment 228\u003cbr\u003e Distributed Environments with Spark - Enabling Data Growth 232\u003cbr\u003e Aggregating Data to Build a Data Mart 237\u003cbr\u003e Visualizing Data with BI Tools 241\u003cbr\u003e 6-2 Data Pipelines with Hadoop 245\u003cbr\u003e Tasking Daily Batch Processing 245\u003cbr\u003e [Task 1] Data Extraction with Embulk 246\u003cbr\u003e [Task 2] Structuring Data with Hive 248\u003cbr\u003e [Task 3] Data Aggregation with Presto 250\u003cbr\u003e 6-3 Automation by Workflow Management Tools 253\u003cbr\u003e Airflow - Script-based Workflow Management 253\u003cbr\u003e Running a Workflow from the Terminal 257\u003cbr\u003e Start the scheduler to run the DAG regularly 260\u003cbr\u003e Controlling the Resources Consumed by Tasks 265\u003cbr\u003e Running Data Pipelines in Hadoop 266 \u003cbr\u003e6-4 Data Pipeline by Cloud Services 268\u003cbr\u003e The Relationship Between Data Analytics and Cloud Services 268\u003cbr\u003e Amazon Web Services 270\u003cbr\u003e Google Cloud Platform 272\u003cbr\u003e Treasure Data 274\u003cbr\u003e 6-5 Summary 279\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eDetailed image\u003c\/b\u003e \u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cdiv\u003e\u003cimg src=\"https:\/\/image.yes24.com\/momo\/TopCate2060\/MidCate004\/205935141.jpg\" border=\"0\" alt=\"Detailed Image 1\"\u003e\u003c\/div\u003e\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eInto the book\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e The term 'big data' began to appear frequently in the media around late 2011 and 2012, around the time when many companies began adopting distributed systems for data processing.\u003cbr\u003e Although data processing by computers had been done before, the term 'big data' began to be used here and there, and the movement to utilize data for business became active.\u003cbr\u003e --- p.3\u003cbr\u003e\u003cbr\u003e The difference between dashboard tools and BI tools is not that strict.\u003cbr\u003e While the former emphasizes the ease of adding new graphs, the latter emphasizes more interactive data exploration. \u003cbr\u003eFor example, if you want to take your time and look at your data slowly, such as clicking on a graph to switch to a detailed view or displaying the raw data that forms the basis for an aggregation, a BI tool is a good choice.\u003cbr\u003e --- p.68\u003cbr\u003e\u003cbr\u003e Hadoop is currently known as a system representing big data, but historically, its development began around 2003 as a distributed file system for 'Nutch', an open source web crawler.\u003cbr\u003e Afterwards, in 2006, it became an independent project and was distributed as Apache Hadoop.\u003cbr\u003e --- p.92\u003cbr\u003e\u003cbr\u003e First, files on object storage are difficult to replace.\u003cbr\u003e Once you write a file, the only way is to replace it in its entirety.\u003cbr\u003e This is fine for things like log files that will not be changed later, but it is not suitable for things that are changed frequently, like databases. \u003cbr\u003eData with a high write frequency should be stored in a separate RDB and snapshotted regularly, or stored in another 'distributed database'.\u003cbr\u003e --- p.161\u003cbr\u003e\u003cbr\u003e Meanwhile, the event data considered in this book, such as messages sent from millions of smartphones, cannot be used directly.\u003cbr\u003e In the previous chapter, we covered the flow of data centered on message flow as a real-time message delivery method.\u003cbr\u003e If batch processing starts with storing the data received in this way in distributed storage, then stream processing is continuing the processing without going through distributed storage.\u003cbr\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e --- p.210 \u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e GOODS SPECIFICS \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eDate of issue:\u003c\/strong\u003e November 5, 2018\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003ePage count, weight, size:\u003c\/strong\u003e 312 pages | 170*225*21mm\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN13:\u003c\/strong\u003e 9791188621439\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN10:\u003c\/strong\u003e 1188621432 \u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cspan\u003e\u003c\/span\u003e\n\n\u003c\/center\u003e\n\n\n\u003c\/center\u003e","brand":"LIBRAIRIE COREENNE","offers":[{"title":"Default Title","offer_id":43893362098218,"sku":"139507","price":37.0,"currency_code":"EUR","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0683\/2750\/5962\/files\/dfed5421a9513b3bf2df3efee0bbbd8a.jpg?v=1765398407","url":"https:\/\/librairie.coreenne.fr\/en\/products\/139507","provider":"LIBRAIRIE COREENNE","version":"1.0","type":"link"}