Skip to product information
Think with data and lead with data in the era of AI upheaval.
In the age of disruptive AI, think with data and lead with data.
Description
Book Introduction
Is your organization making good use of all that dusty data?
You can't discuss technology and business in the AI ​​era without knowing data and statistics!
In a world full of volatility, let's find the patterns behind it!


It lifts the curtain behind data science and provides “the knowledge and know-how to think, speak, understand, and act critically about data.”
From understanding the tendencies of organizational members to the mathematical principles behind algorithms, this book condenses everything about data and statistics used in practice.
This book will help you master the analytical tools, terminology, and mindset necessary to successfully navigate the data science business, while also providing a deeper understanding of challenging data-related problems.
Through learning, you will be able to think critically about data and analytics, and express your opinions intelligently on all things data-related.

  • You can preview some of the book's contents.
    Preview

index
[Part 1] Your First Journey to Thinking and Leading with Data

Chapter 1: What's the Problem?

_Questions Data Leads Must Ask
___Why is this issue important?
___Who does this problem affect?
___What to do if there is no appropriate data
When will the ___ project end?
___What should I do if I am not satisfied with the results?
Why did the data project fail?
___Customer Awareness
___Things to think about
_Let's focus on what matters
_organize

Chapter 2: What is Data?
_Data vs. Information
___Example dataset
_Data type
How is data collected and formatted?
Observational data vs. experimental data
___Structured data vs. unstructured data
_Basic summary statistics
_organize

Chapter 3: Get Ready for Statistical Thinking
_Let's ask a question
_Everything is subject to change
___Customer Awareness Scenario (Sequel)
___Case Study: Kidney Cancer Incidence
_Probability and Statistics
___Probability vs. Intuition
___Discoveries using statistics
_organize

[Part 2] Attitudes toward data, knowledge of probability and statistics

Chapter 4: Let's Argue with Data

_What would you do if you were me?
___The disaster caused by missing data
_Let's check the source of the data
___Who collected the data
___How was the data collected?
Is the data representative?
Was there any bias in the sampling?
How did you handle ___outliers?
_What is unverified data?
___How did you handle missing values?
___Is the data capable of measuring the concept you are trying to measure?
_Let's argue with all data, regardless of size
_organize

Chapter 5: Let's Explore the Data
Exploratory data analysis of data leads
The need for exploratory thinking
__What questions should I ask?
___Virtual Scenario
_Can data answer your questions?
___Set expectations and think sensibly
___Is this a data value that can be intuitively understood?
___Manage outliers and missing values ​​well
_What relationships do you see in the data?
Let's understand the ___ correlation
___Be careful not to misunderstand the correlation.
___Correlation does not imply causation
_Have you found new exploration opportunities in your data?
_organize

Chapter 6: What is Probability?
_Let's guess
_The Rules of the Game
___Mathematical notation
Conditional probability and independent events
___Probability of occurrence of various events
___two events occurring simultaneously
_Thought Experiments on Probability
3 Checkpoints on ___ Probability
_Be careful when assuming that events are independent of each other.
Don't fall into the gambler's fallacy
_Remember that all probabilities are conditional probabilities.
Let's not change the ___dependence relationship
___Bayes' theorem
_Make sure that the probability is meaningful
___correction
___even if the possibility is slim, events happen
_organize

Chapter 7: Let's Challenge Statistics
_What is statistical inference?
___Leave room for error
___The more data there is, the more evidence there is.
___Let's question the current situation
___Is there any evidence to the contrary?
___Balance judgment errors
_Statistical inference process
_Questions needed to verify statistical analysis results
___In what context are the statistical analysis results derived?
___What is the sample size?
___What is being verified
___What is the null hypothesis?
What is the level of ___ significance?
___How many times have you verified it?
Can you provide a confidence interval?
___Is this a practically meaningful result?
Are you assuming a causal relationship?
_organize

[Part 3] Relearning Machine Learning, Deep Learning, and AI Knowledge through Various Case Studies

Chapter 8: Machine Learning: Finding Hidden Patterns and Groups in Data

_What is unsupervised learning?
_Dimension reduction
___Create a complex variable
_Principal component analysis
___Principal components of exercise ability data
___Principal Component Analysis Summary
___Traps to watch out for
_Cluster analysis
_k-means cluster analysis
___Retail store cluster analysis
___Traps to watch out for
_organize

Chapter 9: Regression Models for Predicting the Future and Explaining Phenomena
_Supervised learning
What does linear regression do?
___least squares regression (not just for its fancy name)
_What we can learn from linear regression
___When more variables are input
_The confusion caused by linear regression
___Missing variables
___Multicollinearity
___data leak
___Extrapolation error
___Most relationships are not linear
___Explain or Predict
Performance of the ___regression model
_Other regression models
_organize

Chapter 10: A classification model that can identify the criteria for judgment
_What is a classification problem?
___Three methods of classification models
___Classification problem setting
_Logistic regression
___Advantages of Logistic Regression
_Decision tree
_Ensemble model
___Random Forest
___Gradient Boosted Tree
___Explanatory power of ensemble models
_Beware of common traps
Applying a model that does not fit the ___ data type
___data leak
___Dataset partitioning for model building and testing
___Choosing appropriate thresholds for decision making
_Misconceptions about accuracy
___Confusion matrix
_organize

Chapter 11: Text Analysis: Identifying themes and emotions contained in text
_Expectations for text analysis
_How to convert text to numbers
___word bag
___N-gram
___word embeddings
_Topic Modeling
_Text classification
___Naive Bayes
___Sentiment Analysis
_Practical issues to consider in text analysis
___Big Tech's technological edge
_organize

Chapter 12: Deep Learning and AI: What Data Leaders Need to Know
_Neural network model
___In what ways are neural networks similar to the human brain?
___Simple neural network model
___How Neural Networks Learn
___A slightly more complex neural network
_Deep Learning Application Cases
___Advantages of Deep Learning
___How computers 'see' images
___Convolutional Neural Network
___Deep learning for language processing and sequential data
_Practical Applications of Deep Learning
Is ___ data sufficient?
___Is the data structured?
What does a ___ neural network look like?
_Perspectives on AI
___Big Tech's Advantageous Position
___Ethical Issues of Deep Learning
_organize

[Part 4] What Data Leads Should Do for Project and Organizational Success

Chapter 13: Failures and Traps Lurking Everywhere

_Data Bias and Strange Phenomena
___survivorship bias
___regression to the mean
___Simpson's Paradox
___Confirmation bias
___Sunk cost fallacy
___algorithm bias
___Other biases
_Typical pitfalls of data projects
___Statistics and Machine Learning Pitfalls
___Project Trap
_organize

Chapter 14: Understanding the Diverse Tendencies of Organizational Members
7 Situations When Communication Breaks Down
___Postmortem visit
___a presentation without substance
___Spread of inaccurate information
___into the swamp
___Reality Check
___seize power
___a braggart..
_Three Attitudes People Have When It Comes to Data
___Data Blindman
___Data pessimist
___Data Lead
_organize

Chapter 15: Towards a Higher Place

Detailed image
Detailed Image 1

Into the book
It's an industry that exploits businesses' fears of missing out by promising certainty in an uncertain world.
We call this the 'data science business'.

--- p.38

The stock market fluctuates daily, political polls change weekly (and sometimes even depending on the polling agency), gasoline prices fluctuate, and even though your blood pressure may spike when you get it checked by a doctor, it's normal when you get it checked by a nurse.
If you measure your commute in seconds, it will vary slightly from day to day depending on various factors such as traffic, weather, your child's school commute, and getting your morning coffee out.
There is volatility in everything in the world.
How comfortable are you with this phenomenon?
--- p.91

In casinos, the bags containing the marbles are carefully designed so that samples can be extracted continuously.
But politicians never know what's in the bag until all the marbles (i.e. the voting results) are revealed on Election Day.

--- p.99

When checking a politician's approval rating, if a survey is conducted only targeting supporters of a specific party, the data collected in this way will be subject to sampling bias.
Only well-designed experiments can reduce the risk of potential sampling bias.

--- p.121

Even a child knows that the odds of a coin flip are 50-50, but the 2016 presidential election proved difficult to predict, even with the entire polling industry analyzing terabytes of data.

--- p.149

A debtor's default is not independent of his neighbor's default, but for a long time Wall Street financiers overlooked this fact.
Both events are fundamentally intertwined with the global economic situation.

--- p.158

The false assumption of independence underestimates the likelihood that all projects will fail in the following year, and consequently overestimates the likelihood that at least one project will succeed.
As the 2008 financial crisis and subsequent economic downturn demonstrated, we must not forget the importance of independent assumptions.

--- p.159

I hope you understand clearly that computers cannot understand language like humans, and that to a computer, language is just numbers.
I think just knowing this fact has tremendous value.
If you simply understand that the process of converting text to numbers removes some of the meaning humans assign to words and sentences, you'll never be fooled by marketing claims that AI can solve all your text-related business problems.

--- p.284

It's important to note that algorithmic bias, no matter how well-intentioned (or neutral), can and does occur everywhere.
No model's predictions can tell us the ultimate truth.
Because all results using models are a product of assumptions.

--- p.321

For data pessimists, their personal experiences are more important than data science, statistics, or machine learning.
So they are cynical about the role of data workers.
They view data as a nuisance but an unavoidable element, preferring intuition.
If they are not satisfied with the result, they are more likely to find flaws and overemphasize details rather than offer constructive criticism.
Let's think about why they are so cynical.
--- p.338

Publisher's Review
| What this book covers |
- Attitude and skills toward data for statistical thinking
- Variability that affects daily life and decision-making processes
- Data literacy skills to provide appropriate opinions on statistics and analysis results in the field.
- The fundamental principles and knowledge behind machine learning, text analysis, deep learning, and AI.
- Traps that are easy to fall into when analyzing and interpreting data
- What a data lead must do to ensure the success of projects and organizations.

| Target audience for this book |
This book is fun to read and provides valuable insights for anyone, including novice data scientists, data analysts, business professionals, AI/machine learning engineers, and corporate executives.
This is especially true for marketing professionals who need to work with data analysts, developers, employees, or researchers who are not yet familiar with data, C-level executives who need deeper knowledge of data to introduce new AI technologies and make decisions, and managers who need to lead data teams or organizations.
This is a must-read for anyone who wants to work in the data field or grow into a data leader.


| Structure of this book |

Part 1: Your First Journey to Thinking and Leading with Data
Part 1 covers how to think from a data lead perspective.
Learn how to critically review data projects your organization is undertaking and ask appropriate questions.
We'll explore the definition of data, the correct use of terminology, and how to view the world from a statistical perspective.

Part 2: Attitudes toward data, knowledge of probability and statistics
Data leads actively participate in important discussions about data.
In Part 2, we'll explore how to argue with data and what questions you need to ask to understand the statistical concepts you encounter in your work.
You will learn the basic statistical and probability concepts needed to understand or address data analysis results.

Part 3: Relearning Machine Learning, Deep Learning, and AI Knowledge through Various Case Studies
Data leaders must understand the fundamental principles of how statistical and machine learning models work.
You will gain an intuitive understanding of unsupervised learning, regression, classification, text analysis, and deep learning.

Part 4: What Data Leads Do to Ensure Project and Organizational Success
Data leads should be aware of common mistakes and pitfalls when working with data.
We explore the technical pitfalls that lead organizations and projects to failure, and learn about the people involved in data projects and their tendencies.
Finally, we will provide guidance on how to succeed as a data leader.
[Author's Note]

We believe many people want to learn about data properly, but don't know where to start.
The numerous data science and statistics books already published cover a very diverse spectrum.
On one side of the spectrum, there are numerous non-technical books that extol the benefits and hopeful outlook of data utilization.
There are some good books among them.
However, no matter how good a book is, it is often a book that deals with business that is biased towards the present.
Most of the books are written by journalists to highlight the dramatic aspects of the rise of data.

These books explain how specific business problems are solved through the lens of data. Terms like AI and machine learning often appear.
Here, I hope there is no misunderstanding.
Books like this have helped raise people's awareness of how to use data.
However, it does not delve deeply into the specific tasks to be performed, but simply focuses on the big picture problems and solutions.

At the other end of the spectrum are high-level technical books.
Most of them are heavy books with over 500 pages, so they are not only physically demanding but also mentally taxing.

There are mountains of books at both ends of the spectrum.
The communication gap is not easily narrowed because most people only read business or technical books.
Fortunately, there are some excellent books between these two extremes.


We'd like to add our book to this list.
The book you are reading can be read by anyone without any burden, even without a computer or notepad nearby.
If you enjoyed our book, I highly recommend reading at least one of the two books mentioned above to further solidify your understanding of data.
You won't regret it.

Our authors really like this book.
If this book motivates you to learn about data and data analysis and sparks a desire to learn more, then I consider that a success.

[Translator's Note]

I've read, studied, and sometimes translated many books on data, but this book is so great that I wish I had written it myself instead of translating it.
When I first received the original book and skimmed through the chapter titles, I thought, "Isn't this content too easy?" However, from the moment I started reading it in earnest, savoring each sentence in preparation for the translation, until the very end of the last chapter, I couldn't help but be impressed by the authors' efforts to write in line with the book's planning intentions, and their deep expertise in data analysis and statistics.

It is often said that “the easiest thing to write is the hardest.”
I've always rationally agreed with this statement, but I've rarely experienced a concrete example. After reading this book, I felt like I'd finally encountered a true example of what it means.
'Easy to write' means that the writer has a perfect grasp of the core and logic of the content, which allows for writing that is easy, clear, and logical.

This book covers just the right depth and breadth of data analysis and statistics, which can be challenging.
It's a great introductory book for those considering a career in this field, but it's also incredibly helpful for anyone who doesn't need to delve too deeply into the technical details but wants to build up their knowledge to a level where they can communicate with data analysts.
It's a truly remarkable book that delicately walks the line between general education and a full-fledged technical book.

Especially in an era like today, when AI is rapidly becoming popular, it is necessary to focus on the 'essence' of data, the raw material that powers AI.
While countless books and articles about AI abound today, the most accurate way to understand AI is to trace the evolution of "data-driven statistical thinking" into AI.
In that sense, I hope this book can serve as a first textbook for the general public living in the AI ​​era.

The technical aspects of the book are something I'm already familiar with, and since they don't go into too much detail, I was able to easily understand the authors' arguments and messages in the original book. However, the problem was the process of translating it into Korean.
The content covered in each sentence and paragraph is dense and the meaning is compressed, so the sentence itself is easy, but it took a lot of thought and time to translate the exact meaning and subtle nuances of the original text into Korean sentences.
Although it is an old saying, I had no choice but to translate it 'one by one', putting in a lot of time and effort.

I must confess that I have always had a strong desire to write a good book that strikes the perfect balance between educational and technical.
However, while translating this book, I felt a sense of disappointment that a similar book had already been published, but at the same time, I felt a sense of joy at having discovered such a great book and being given the opportunity to translate it.
It is such an excellent book, and I can confidently recommend it to many people.
- Choi Jae-won

Having lived as a materials engineer for several decades, I have been busy researching and analyzing various materials engineering phenomena, even during my degree course.
After receiving my degree, I started working in the semiconductor industry. In addition to the materials engineering perspective I had previously covered, I was exposed to statistical concepts such as various quality control techniques and model interpretation for reliability analysis.
It was essential for companies to improve the quality and lifespan of their products to make a profit.

But there was still a thought lingering in my head.
If we fully understand the phenomena dealt with in materials engineering, it seemed that such statistical approaches could be minimized and perhaps even eliminated.
Looking back now, I think perhaps it was more that I wanted to completely ignore statistical approaches and applications.
Most of the science and engineering I've been focusing on, including materials engineering, has been about exploring causal relationships to clearly identify cause and effect.
Then, the AI ​​era arrived and began to be applied to all fields, including semiconductors.
This made me wonder what the difference was between the classical data concepts in statistics and the data handled by AI.
Out of this vague curiosity, I searched through countless papers and books, and even wandered the ocean of the Internet.

For ordinary researchers like me, isn't there a book like "a pearl in the mud" that doesn't just talk about rosy futures, without coding or complex statistical formulas, but instead gets straight to the point? Are there authors who write for people with similar curiosity? Indeed, this connection did exist! It was the original book, "Becoming a Data Head."
This pearl, which I found with great difficulty, was an English book, but it was so enjoyable to read that I still vividly remember the feeling.
The authors of this book were storytellers who, as if they were discussing everything I had ever wondered about in a chatroom, were able to smoothly unravel it. As I read the book, I felt as if decades of fundamental questions about statistics and data had been answered in one fell swoop.

How many things in the universe, including technical challenges, can we truly understand causal relationships? That's why it was necessary to start with statistics and understand the world of data opened up by deep learning and AI.
Another striking aspect of this book is that it offers a variety of metaphors that illustrate how not only ordinary engineers and researchers, but also business executives and managers should view and utilize data for the success of their companies.

I hope that readers from various fields will enjoy the joy of data enlightenment that this book brings, and I conclude this article with that hope.
- Jang Jin-wook [Author's Note]

We believe many people want to learn about data properly, but don't know where to start.
The numerous data science and statistics books already published cover a very diverse spectrum.
On one side of the spectrum, there are numerous non-technical books that extol the benefits and hopeful outlook of data utilization.
There are some good books among them.
However, no matter how good a book is, it is often a book that deals with business that is biased towards the present.
Most of the books are written by journalists to highlight the dramatic aspects of the rise of data.

These books explain how specific business problems are solved through the lens of data. Terms like AI and machine learning often appear.
Here, I hope there is no misunderstanding.
Books like this have helped raise people's awareness of how to use data.
However, it does not delve deeply into the specific tasks to be performed, but simply focuses on the big picture problems and solutions.

At the other end of the spectrum are high-level technical books.
Most of them are heavy books with over 500 pages, so they are not only physically demanding but also mentally taxing.

There are mountains of books at both ends of the spectrum.
The communication gap is not easily narrowed because most people only read business or technical books.
Fortunately, there are some excellent books between these two extremes.


We'd like to add our book to this list.
The book you are reading can be read by anyone without any burden, even without a computer or notepad nearby.
If you enjoyed our book, I highly recommend reading at least one of the two books mentioned above to further solidify your understanding of data.
You won't regret it.

Our authors really like this book.
If this book motivates you to learn about data and data analysis and sparks a desire to learn more, then I consider that a success.

[Translator's Note]

I've read, studied, and sometimes translated many books on data, but this book is so great that I wish I had written it myself instead of translating it.
When I first received the original book and skimmed through the chapter titles, I thought, "Isn't this content too easy?" However, from the moment I started reading it in earnest, savoring each sentence in preparation for the translation, until the very end of the last chapter, I couldn't help but be impressed by the authors' efforts to write in line with the book's planning intentions, and their deep expertise in data analysis and statistics.

It is often said that “the easiest thing to write is the hardest.”
I've always rationally agreed with this statement, but I've rarely experienced a concrete example. After reading this book, I felt like I'd finally encountered a true example of what it means.
'Easy to write' means that the writer has a perfect grasp of the core and logic of the content, which allows for writing that is easy, clear, and logical.

This book covers just the right depth and breadth of data analysis and statistics, which can be challenging.
It's a great introductory book for those considering a career in this field, but it's also incredibly helpful for anyone who doesn't need to delve too deeply into the technical details but wants to build up their knowledge to a level where they can communicate with data analysts.
It's a truly remarkable book that delicately walks the line between general education and a full-fledged technical book.

Especially in an era like today, when AI is rapidly becoming popular, it is necessary to focus on the 'essence' of data, the raw material that powers AI.
While countless books and articles about AI abound today, the most accurate way to understand AI is to trace the evolution of "data-driven statistical thinking" into AI.
In that sense, I hope this book can serve as a first textbook for the general public living in the AI ​​era.

The technical aspects of the book are something I'm already familiar with, and since they don't go into too much detail, I was able to easily understand the authors' arguments and messages in the original book. However, the problem was the process of translating it into Korean.
The content covered in each sentence and paragraph is dense and the meaning is compressed, so the sentence itself is easy, but it took a lot of thought and time to translate the exact meaning and subtle nuances of the original text into Korean sentences.
Although it is an old saying, I had no choice but to translate it 'one by one', putting in a lot of time and effort.

I must confess that I have always had a strong desire to write a good book that strikes the perfect balance between educational and technical.
However, while translating this book, I felt a sense of disappointment that a similar book had already been published, but at the same time, I felt a sense of joy at having discovered such a great book and being given the opportunity to translate it.
It is such an excellent book, and I can confidently recommend it to many people.
- Choi Jae-won

Having lived as a materials engineer for several decades, I have been busy researching and analyzing various materials engineering phenomena, even during my degree course.
After receiving my degree, I started working in the semiconductor industry. In addition to the materials engineering perspective I had previously covered, I was exposed to statistical concepts such as various quality control techniques and model interpretation for reliability analysis.
It was essential for companies to improve the quality and lifespan of their products to make a profit.

But there was still a thought lingering in my head.
If we fully understand the phenomena dealt with in materials engineering, it seemed that such statistical approaches could be minimized and perhaps even eliminated.
Looking back now, I think perhaps it was more that I wanted to completely ignore statistical approaches and applications.
Most of the science and engineering I've been focusing on, including materials engineering, has been about exploring causal relationships to clearly identify cause and effect.
Then, the AI ​​era arrived and began to be applied to all fields, including semiconductors.
This made me wonder what the difference was between the classical data concepts in statistics and the data handled by AI.
Out of this vague curiosity, I searched through countless papers and books, and even wandered the ocean of the Internet.

For ordinary researchers like me, isn't there a book like "a pearl in the mud" that doesn't just talk about rosy futures, without coding or complex statistical formulas, but instead gets straight to the point? Are there authors who write for people with similar curiosity? Indeed, this connection did exist! It was the original book, "Becoming a Data Head."
This pearl, which I found with great difficulty, was an English book, but it was so enjoyable to read that I still vividly remember the feeling.
The authors of this book were storytellers who, as if they were discussing everything I had ever wondered about in a chatroom, were able to smoothly unravel it. As I read the book, I felt as if decades of fundamental questions about statistics and data had been answered in one fell swoop.

How many things in the universe, including technical challenges, can we truly understand causal relationships? That's why it was necessary to start with statistics and understand the world of data opened up by deep learning and AI.
Another striking aspect of this book is that it offers a variety of metaphors that illustrate how not only ordinary engineers and researchers, but also business executives and managers should view and utilize data for the success of their companies.

I hope that readers from various fields will enjoy the joy of data enlightenment that this book brings, and I conclude this article with that hope.
- Jang Jin-wook
GOODS SPECIFICS
- Date of issue: May 3, 2024
- Page count, weight, size: 368 pages | 544g | 152*224*19mm
- ISBN13: 9791189909628
- ISBN10: 1189909626

You may also like

카테고리