
Real-world datasets for machine learning
Description
Book Introduction
Balancing privacy and broad data usage.
Building and testing machine learning models requires large and diverse types of data.
However, most datasets are not widely available due to privacy concerns.
This book introduces practical synthetic data techniques for creating new data from real data.
Synthetic data is amenable to secondary analysis and can be used for a variety of purposes, including data research, understanding customer behavior, and developing new products.
This book provides methods for synthesizing real-world data for use in a variety of industries and addresses privacy concerns.
You will also learn the principles and steps for generating synthetic data from real datasets.
Furthermore, you will learn how synthetic data can reduce the time it takes to develop products or solutions.
Building and testing machine learning models requires large and diverse types of data.
However, most datasets are not widely available due to privacy concerns.
This book introduces practical synthetic data techniques for creating new data from real data.
Synthetic data is amenable to secondary analysis and can be used for a variety of purposes, including data research, understanding customer behavior, and developing new products.
This book provides methods for synthesizing real-world data for use in a variety of industries and addresses privacy concerns.
You will also learn the principles and steps for generating synthetic data from real datasets.
Furthermore, you will learn how synthetic data can reduce the time it takes to develop products or solutions.
- You can preview some of the book's contents.
Preview
index
CHAPTER 1: INTRODUCING SYNTHETIC DATA CREATION
1.1 Defining synthetic data
1.2 Benefits of Synthetic Data
1.3 Use Cases of Synthetic Data
1.4 Summary
CHAPTER 2 DATA SYNTHESIS
2.1 Synthesis period
2.2 Identification Possibility Spectrum
2.3 Compromise in PET Selection to Enable Data Access
2.4 Data Synthesis Project
2.5 Data Synthesis Pipeline
2.6 Managing Synthetic Programs
2.7 Summary
CHAPTER 3 Getting Started: Distribution Fitting
3.1 Data Frame
3.2 Data distribution types
3.3 Fitting the distribution to real data
3.4 Generating synthetic data from distributions
3.5 Summary
CHAPTER 4 Evaluating the Utility of Synthetic Data
4.1 Synthetic Data Utility Framework: Analytical Replication
4.2 Synthetic Data Utility Framework: Utility Metrics
4.3 Summary
CHAPTER 5 Data Synthesis Methods
5.1 Synthetic Data Generation Theory
5.2 Generating Real Synthetic Data
5.3 Hybrid Synthetic Data
5.4 Machine Learning Methods
5.5 Deep Learning Methods
5.6 Sequence Synthesis
5.7 Summary
CHAPTER 6: Identifying Synthetic Data
6.1 Exposure Types
6.2 The Impact of Privacy Laws on the Creation and Use of Synthetic Data
6.3 Summary
CHAPTER 7 Real Data Synthesis
7.1 Managing Data Complexity
7.2 Configuring data synthesis
7.3 Conclusion
1.1 Defining synthetic data
1.2 Benefits of Synthetic Data
1.3 Use Cases of Synthetic Data
1.4 Summary
CHAPTER 2 DATA SYNTHESIS
2.1 Synthesis period
2.2 Identification Possibility Spectrum
2.3 Compromise in PET Selection to Enable Data Access
2.4 Data Synthesis Project
2.5 Data Synthesis Pipeline
2.6 Managing Synthetic Programs
2.7 Summary
CHAPTER 3 Getting Started: Distribution Fitting
3.1 Data Frame
3.2 Data distribution types
3.3 Fitting the distribution to real data
3.4 Generating synthetic data from distributions
3.5 Summary
CHAPTER 4 Evaluating the Utility of Synthetic Data
4.1 Synthetic Data Utility Framework: Analytical Replication
4.2 Synthetic Data Utility Framework: Utility Metrics
4.3 Summary
CHAPTER 5 Data Synthesis Methods
5.1 Synthetic Data Generation Theory
5.2 Generating Real Synthetic Data
5.3 Hybrid Synthetic Data
5.4 Machine Learning Methods
5.5 Deep Learning Methods
5.6 Sequence Synthesis
5.7 Summary
CHAPTER 6: Identifying Synthetic Data
6.1 Exposure Types
6.2 The Impact of Privacy Laws on the Creation and Use of Synthetic Data
6.3 Summary
CHAPTER 7 Real Data Synthesis
7.1 Managing Data Complexity
7.2 Configuring data synthesis
7.3 Conclusion
Publisher's Review
- Generating synthetic data using multivariate normal distribution
- Fitting distributions of various fitness metrics
- Replicating the structure of the original data
- Modeling data with complex relationships
- Establishing methods and measurement criteria for evaluating data utility.
- Analyze real data to replicate synthetic data.
- Assessing the privacy and identity exposure of synthetic data
Synthetic data has gained significant attention and societal interest in the past few years, driven by two key concerns:
First, there is the demand for massive amounts of data to train and build artificial intelligence and machine learning (AIML) models.
Second, recent work demonstrates an effective method for generating high-quality synthetic data.
This has led to a recognition that synthetic data can be quite effective at solving some challenging problems, especially within the AIML community.
As a result, not only companies like NVIDIA, IBM, and Alphabet, but also government agencies like the U.S. Census Bureau have adopted various types of data synthesis methodologies to support model building, application development, and data distribution.
Chapter 1: Explains synthetic data and its benefits.
Artificial intelligence and machine learning (AIML) projects are being used across a wide range of industries, and we've included a few sample use cases to give you a taste of their breadth.
Chapter 2: We present a decision-making framework to help you set goals for data synthesis and determine when it best suits your business priorities compared to other methods.
Chapter 3: Covers distribution modeling, the first step in the data synthesis process.
We outline how to fit non-standard data distributions to machine learning models.
Chapter 4: Describes a data utility framework that can be used with synthetic data.
We explore data synthesis optimization, data synthesis approaches, and understanding the results of synthetic data.
Chapter 5: Creating synthetic data using basic concepts.
It starts with some basic approaches and progresses to more complex ones, covering techniques from beginner to advanced.
Chapter 6: First, we define the types of exposures that data synthesis aims to protect.
We examine how key privacy regulations in the US and EU address synthetic data and suggest ways to begin privacy-preserving analytics.
Chapter 7: Building on our experience teaching synthetic datasets and synthetic data generation techniques, we present practical considerations that will be helpful when processing real data.
It not only highlights challenging tasks, but also suggests ways to solve them.
- Fitting distributions of various fitness metrics
- Replicating the structure of the original data
- Modeling data with complex relationships
- Establishing methods and measurement criteria for evaluating data utility.
- Analyze real data to replicate synthetic data.
- Assessing the privacy and identity exposure of synthetic data
Synthetic data has gained significant attention and societal interest in the past few years, driven by two key concerns:
First, there is the demand for massive amounts of data to train and build artificial intelligence and machine learning (AIML) models.
Second, recent work demonstrates an effective method for generating high-quality synthetic data.
This has led to a recognition that synthetic data can be quite effective at solving some challenging problems, especially within the AIML community.
As a result, not only companies like NVIDIA, IBM, and Alphabet, but also government agencies like the U.S. Census Bureau have adopted various types of data synthesis methodologies to support model building, application development, and data distribution.
Chapter 1: Explains synthetic data and its benefits.
Artificial intelligence and machine learning (AIML) projects are being used across a wide range of industries, and we've included a few sample use cases to give you a taste of their breadth.
Chapter 2: We present a decision-making framework to help you set goals for data synthesis and determine when it best suits your business priorities compared to other methods.
Chapter 3: Covers distribution modeling, the first step in the data synthesis process.
We outline how to fit non-standard data distributions to machine learning models.
Chapter 4: Describes a data utility framework that can be used with synthetic data.
We explore data synthesis optimization, data synthesis approaches, and understanding the results of synthetic data.
Chapter 5: Creating synthetic data using basic concepts.
It starts with some basic approaches and progresses to more complex ones, covering techniques from beginner to advanced.
Chapter 6: First, we define the types of exposures that data synthesis aims to protect.
We examine how key privacy regulations in the US and EU address synthetic data and suggest ways to begin privacy-preserving analytics.
Chapter 7: Building on our experience teaching synthetic datasets and synthetic data generation techniques, we present practical considerations that will be helpful when processing real data.
It not only highlights challenging tasks, but also suggests ways to solve them.
GOODS SPECIFICS
- Date of issue: January 4, 2021
- Page count, weight, size: 172 pages | 183*235*20mm
- ISBN13: 9791162243749
- ISBN10: 1162243740
You may also like
카테고리
korean
korean