
Language models and generative AI starting with Transformers
Description
Book Introduction
Analyze pre-trained language models, giant language models, and small giant language models starting with Transformers!
A detailed analysis of the rapid changes in artificial intelligence technology since ChatGPT!
Analyzing the diverse language models used by tech giants competing against OpenAI!
This book explains the basic principles of transformers, the development of language models based on them, and the era of generative artificial intelligence in an easy and systematic way.
In particular, rather than simply listing the core principles and technologies of artificial intelligence, it is organized so that they can be intuitively understood through 100 topics and illustrations divided by era.
For readers who have been hesitant about learning artificial intelligence due to complex formulas and algorithms, this book will serve as a stepping stone to take the next step.
A detailed analysis of the rapid changes in artificial intelligence technology since ChatGPT!
Analyzing the diverse language models used by tech giants competing against OpenAI!
This book explains the basic principles of transformers, the development of language models based on them, and the era of generative artificial intelligence in an easy and systematic way.
In particular, rather than simply listing the core principles and technologies of artificial intelligence, it is organized so that they can be intuitively understood through 100 topics and illustrations divided by era.
For readers who have been hesitant about learning artificial intelligence due to complex formulas and algorithms, this book will serve as a stepping stone to take the next step.
- You can preview some of the book's contents.
Preview
index
Chapter 1: Transformers: Innovation in Artificial Intelligence
01 Artificial Intelligence Changed by Transformers
02 Conceptual Understanding of Transformers
03 Understanding Language Models
04 Language Model and Data
05 Changes in language model implementation methods
06 Pre-trained language model
07 Giant Language Model
08-based models and generative artificial intelligence
Chapter 2: Structure and Analysis of Transformers
01 Deep Learning and Transformers
02 Transformer of artificial intelligence models
03 Transformer Structure
04 Transformer Input - Tokenization
05 Transformer Input - Setting Token Relationships
06 One-Hot Encoding and Token Embedding (CBOW)
07 Token Embedding and Token Embedding Dimensions
08 Positional Encoding - Token Location and Access
09 Positional Encoding - Token Embedding Vector
10 Transformer Encoder
11 Input and query of multihead attention
12 Key and dot product attention of multihead attention
13 The Value of Multihead Attention and Attention Heads
14 Regularization and Feedforward Neural Networks
15 Self-Attention
16 Encoder iterations and hyperparameters
17 Transformer's decoder
18 Combination of encoder and decoder of transformer
19 Transformer output
Chapter 3 Pre-trained language models
01 Overview of Pre-trained Language Models
02 Pre-trained language model approach
03 Various natural language processing tasks
04 Typical Natural Language Processing Tasks
05 BERT's Structure and Features
06 OpenAI's GPT Structure
07 Concept and Structure of GPT-2
08 RoBERTa's performance and structure
09 ALBERT's approach and structure
10 DistilBERT's approach and performance
11. Concept and Features of MobileBERT
12 Concept and Features of SpanBERT
13 Concept and Application of ELECTRA
14 Concept and Function of DeBERTa
15 The appearance and performance of TransformerXL
16 XLNet Concept and Performance
17 Concept and Features of BART
18 CTRL concept and features
Structure and performance of the 19 T5
20 HuggingFace and Transformers
Chapter 4: Large Language Models
01 Overview of the Giant Language Model
02 Scale of the large language model
03 Structure of a large language model
04 Characteristics of large language models and retraining methods
05 Limitations of Large Language Models
06 Deep Learning from a Computational Perspective
07 Matrix multiplication operation
08 Computational Amount and Effectiveness of Large Language Models
09 Utilizing large language models
10 Structure and Performance of GPT-3
11 Structure and Features of LaMDA
Concept and Performance of 12 MT NLG
13 The emergence and approach of Gopher
14 InstructGPT's Approach and Features
15 Background and Features of PanGu Alpha
16 Concept and Features of PaLM
17 OPT 175B Introduction and Features
18 BLOOM's background and features
19 Concept and Features of HyperCLOVA
20-Size Competition and the Emergence of ChatGPT
Chapter 5: ChatGPT and Generative Artificial Intelligence
01 ChatGPT's Success and Changes
02 ChatGPT's effectiveness and computational complexity
03 Artificial Intelligence after ChatGPT
04 Basic Approaches to Alleviating Hallucinogenic Effects
05 Explainable AI and Large Language Models
06 Basic approach to model lightweighting
07 Model Lightweight and Hardware
08 Basic Approach to Transformers
09 Concept and Structure of RWKV
10 Approaches and Features of Retentive Networks
11 Moves by Giant Tech Companies
12 Background and Features of GPT-4
13 Features of GPT-4 Turbo and GPT-4o
14 Features of GPT-o1 and GPT-5
15 OpenAI's GPT Store
16 Features and Functions of Bard
17 The size and features of PaLM 2
18 Gemini's appearance and characteristics
19 The appearance and features of Copilot
20 Claude's Development and Features
21 The emergence of small-scale giant language models
22 Characteristics and Structure of LLaMA
23 Performance of LLaMA 2 and LLaMA 3
24 Gemma Features and Functions
25 The emergence and features of Mistral AI
26 Features and functions of phi
27 Korea's sLLM and Solar
28 Conversational AI and sLLM Utilization Tools
29 Technological Future Prospects for Artificial Intelligence
30 Future Prospects for Artificial Intelligence Platforms
31 Regulatory Outlook for Artificial Intelligence
32 Policy Outlook for Artificial Intelligence
33 Future Prospects for General Artificial Intelligence
01 Artificial Intelligence Changed by Transformers
02 Conceptual Understanding of Transformers
03 Understanding Language Models
04 Language Model and Data
05 Changes in language model implementation methods
06 Pre-trained language model
07 Giant Language Model
08-based models and generative artificial intelligence
Chapter 2: Structure and Analysis of Transformers
01 Deep Learning and Transformers
02 Transformer of artificial intelligence models
03 Transformer Structure
04 Transformer Input - Tokenization
05 Transformer Input - Setting Token Relationships
06 One-Hot Encoding and Token Embedding (CBOW)
07 Token Embedding and Token Embedding Dimensions
08 Positional Encoding - Token Location and Access
09 Positional Encoding - Token Embedding Vector
10 Transformer Encoder
11 Input and query of multihead attention
12 Key and dot product attention of multihead attention
13 The Value of Multihead Attention and Attention Heads
14 Regularization and Feedforward Neural Networks
15 Self-Attention
16 Encoder iterations and hyperparameters
17 Transformer's decoder
18 Combination of encoder and decoder of transformer
19 Transformer output
Chapter 3 Pre-trained language models
01 Overview of Pre-trained Language Models
02 Pre-trained language model approach
03 Various natural language processing tasks
04 Typical Natural Language Processing Tasks
05 BERT's Structure and Features
06 OpenAI's GPT Structure
07 Concept and Structure of GPT-2
08 RoBERTa's performance and structure
09 ALBERT's approach and structure
10 DistilBERT's approach and performance
11. Concept and Features of MobileBERT
12 Concept and Features of SpanBERT
13 Concept and Application of ELECTRA
14 Concept and Function of DeBERTa
15 The appearance and performance of TransformerXL
16 XLNet Concept and Performance
17 Concept and Features of BART
18 CTRL concept and features
Structure and performance of the 19 T5
20 HuggingFace and Transformers
Chapter 4: Large Language Models
01 Overview of the Giant Language Model
02 Scale of the large language model
03 Structure of a large language model
04 Characteristics of large language models and retraining methods
05 Limitations of Large Language Models
06 Deep Learning from a Computational Perspective
07 Matrix multiplication operation
08 Computational Amount and Effectiveness of Large Language Models
09 Utilizing large language models
10 Structure and Performance of GPT-3
11 Structure and Features of LaMDA
Concept and Performance of 12 MT NLG
13 The emergence and approach of Gopher
14 InstructGPT's Approach and Features
15 Background and Features of PanGu Alpha
16 Concept and Features of PaLM
17 OPT 175B Introduction and Features
18 BLOOM's background and features
19 Concept and Features of HyperCLOVA
20-Size Competition and the Emergence of ChatGPT
Chapter 5: ChatGPT and Generative Artificial Intelligence
01 ChatGPT's Success and Changes
02 ChatGPT's effectiveness and computational complexity
03 Artificial Intelligence after ChatGPT
04 Basic Approaches to Alleviating Hallucinogenic Effects
05 Explainable AI and Large Language Models
06 Basic approach to model lightweighting
07 Model Lightweight and Hardware
08 Basic Approach to Transformers
09 Concept and Structure of RWKV
10 Approaches and Features of Retentive Networks
11 Moves by Giant Tech Companies
12 Background and Features of GPT-4
13 Features of GPT-4 Turbo and GPT-4o
14 Features of GPT-o1 and GPT-5
15 OpenAI's GPT Store
16 Features and Functions of Bard
17 The size and features of PaLM 2
18 Gemini's appearance and characteristics
19 The appearance and features of Copilot
20 Claude's Development and Features
21 The emergence of small-scale giant language models
22 Characteristics and Structure of LLaMA
23 Performance of LLaMA 2 and LLaMA 3
24 Gemma Features and Functions
25 The emergence and features of Mistral AI
26 Features and functions of phi
27 Korea's sLLM and Solar
28 Conversational AI and sLLM Utilization Tools
29 Technological Future Prospects for Artificial Intelligence
30 Future Prospects for Artificial Intelligence Platforms
31 Regulatory Outlook for Artificial Intelligence
32 Policy Outlook for Artificial Intelligence
33 Future Prospects for General Artificial Intelligence
Detailed image

Publisher's Review
This book goes beyond simply explaining Transformers.
Since the advent of Transformers, artificial intelligence has entered a new era of generative AI, and the barriers to entry for development have been significantly lowered through platforms and open source.
In particular, we analyze generative AI from the perspective of language models to explain the competitive landscape of global big tech companies.
Additionally, we provide an opportunity to utilize real-world artificial intelligence by providing videos of pre-trained language models and sLLM exercises that anyone can easily run on a laptop.
I hope this book will serve as a valuable guide for students, researchers, and practitioners exploring the world of Transformers and generative AI. The pace of AI technology development is astonishingly rapid.
But at the heart of that innovation has always been, and will continue to be, Transformers.
A word from the author
The advent of Transformers has shifted the paradigm of natural language processing (NLP) and generative AI, and models like GPT, BERT, and DALL-E, which were born from them, are now revolutionizing our daily lives and industries.
In this way, Transformer represents a paradigm shift beyond a single deep learning model.
When recurrent neural networks (RNNs) and long short-term memory (LSTMs) in the past showed limitations by failing to overcome temporal dependence, Transformers overcame these limitations through the paper “Attention is all you need.”
As a result, Transformers have become a new standard in natural language processing, being used in a variety of fields, including machine translation, language understanding, conversational systems, and text generation.
Subsequent extensions of pre-trained language models (PLMs) and large language models (LLMs) have rewritten the history of artificial intelligence, proving the potential of transformers.
The future of artificial intelligence depends on how we learn and utilize it.
Now, let's begin that journey together.
Since the advent of Transformers, artificial intelligence has entered a new era of generative AI, and the barriers to entry for development have been significantly lowered through platforms and open source.
In particular, we analyze generative AI from the perspective of language models to explain the competitive landscape of global big tech companies.
Additionally, we provide an opportunity to utilize real-world artificial intelligence by providing videos of pre-trained language models and sLLM exercises that anyone can easily run on a laptop.
I hope this book will serve as a valuable guide for students, researchers, and practitioners exploring the world of Transformers and generative AI. The pace of AI technology development is astonishingly rapid.
But at the heart of that innovation has always been, and will continue to be, Transformers.
A word from the author
The advent of Transformers has shifted the paradigm of natural language processing (NLP) and generative AI, and models like GPT, BERT, and DALL-E, which were born from them, are now revolutionizing our daily lives and industries.
In this way, Transformer represents a paradigm shift beyond a single deep learning model.
When recurrent neural networks (RNNs) and long short-term memory (LSTMs) in the past showed limitations by failing to overcome temporal dependence, Transformers overcame these limitations through the paper “Attention is all you need.”
As a result, Transformers have become a new standard in natural language processing, being used in a variety of fields, including machine translation, language understanding, conversational systems, and text generation.
Subsequent extensions of pre-trained language models (PLMs) and large language models (LLMs) have rewritten the history of artificial intelligence, proving the potential of transformers.
The future of artificial intelligence depends on how we learn and utilize it.
Now, let's begin that journey together.
GOODS SPECIFICS
- Date of issue: January 13, 2025
- Page count, weight, size: 220 pages | 187*240*20mm
- ISBN13: 9791198685339
- ISBN10: 1198685336
You may also like
카테고리
korean
korean