Skip to product information
reinforcement learning
reinforcement learning
Description
Book Introduction
Understand key concepts and algorithms
A textbook on reinforcement learning implemented through practice

"Reinforcement Learning" is aimed at readers who want to systematically learn reinforcement learning from the basics to its applications.
First, we introduce a new idea for solving the problem and explain the basic principles that support it.
Next, we present a formula-based algorithm and organize the content so that students can implement the algorithm in a real-world environment and evaluate its learning performance by practicing programs using TensorFlow and Gymnasium.
It also covers everything from traditional dynamic programming to the latest algorithms such as DQN, PPO, and SAC, and faithfully reflects advanced application areas such as autonomous driving, intelligent robots, and AGI, as well as the latest research trends.

※ This book was developed as a textbook for university lectures, so it does not provide answers to practice problems.
  • You can preview some of the book's contents.
    Preview

index
CHAPTER 01 Introduction
1 What is reinforcement learning?
2 Machine Learning = Supervised Learning + Unsupervised Learning + Reinforcement Learning
3 Success Stories and Applications
4 Brief History
5 Readings
Practice problems

CHAPTER 02 Establishing the Basics of Reinforcement Learning
1 Agents that interact with the environment
2 MDP Programming with Python
3 Comparison of expected gains between random and optimal policies
4 Policy and Value Functions
5 Understanding the Difficulty and Approaches of Reinforcement Learning
Practice problems

CHAPTER 03 Dynamic Programming
1 Principle
2 Bellman equation and policy iteration algorithm
3 Bellman Optimality Equation and Value Iteration Algorithm
4 Dynamic Programming of Stochastic Tasks
5 Characteristics and Limitations of Dynamic Programming
Practice problems

CHAPTER 04 Monte Carlo Method
1 episode generator
2 Balance of Exploration and Exploration
3 Policy Evaluation Using the Monte Carlo Method
4 Policy learning using the Monte Carlo method
5 Performance Improvement Techniques
Practice problems

CHAPTER 05 Time-Differential Learning
1 Principle
2 Policy Evaluation
3 Sarsa
4 Q-Learning
Learning Blackjack with Q-Learning
6 Training CartPole with Q-Learning
7 Performance Improvement Techniques
8 Expanding the perspective
Practice problems

CHAPTER 06 Approximation Methods Using Neural Networks
1 Neural Network Basics
2 Neural network implementation
Q-learning using 3 neural networks
4 Neural Network-Based Q-Learning Implementation: CartPole Task
5 Neural Network-Based Q-Learning Implementation: Blackjack Task
6 Controversies and New Paths for Neural Networks
Practice problems

CHAPTER 07 Deep Learning Methods
1 A major shift towards deep learning
2 DQNs
3 Replay Memory
4 Deep Learning Basics
5 Atari gaming environment
6 Pong Atari game using DQN
7 Additional remarks
Practice problems

CHAPTER 08 Policy Gradient Method
1 Policy-centered learning
2 REINFORCE algorithm
3 REINFORCE Programming: Discrete Tasks
Policy gradient for 4 consecutive tasks
5 REINFORCE Programming: Continuous Tasks
6 Speeding Up Python's Array Operations
Practice problems

CHAPTER 09 The Activist-Critic Method
1. Collaboration between activists and critics
2 Bias-Variance Tradeoff
3 Profit function
4 A2C and A3C
5 A2C Programming: Discrete Tasks
6 A2C Programming: Continuous Tasks
Practice problems

CHAPTER 10 TRUST REGION METHOD
1. Improvement of the single-note policy
2 TRPO algorithm
3 PPO algorithm
4 Improving the efficiency of PPO
5 PPO Programming: Sequential Tasks
Practice problems

CHAPTER 11 Combining Policy Optimization and DQN
1 Motivation and Development
2 DDPG learning algorithm
3 Programming Practice: Learning Hopper Tasks Using DDPG
4 TD3 learning algorithm
5 Programming Practice: Learning the Hopper Task Using TD3
6 SAC learning algorithms
7 Programming Practice: Learning Hopper Tasks with SAC
8 Benchmarking Analysis
Practice problems

CHAPTER 12 Learning by Mimicry
1. Idea and Development
2. Duplicate actions
3. Inverse reinforcement learning
4. Adversarial mimicry learning
5. Observation Mimicry
Practice problems

CHAPTER 13 ADVANCED APPLICATIONS
1 Advanced Products Built with Reinforcement Learning: Opportunities and Challenges
2 video games
3 board games
4 Large-Scale Language Models
5 Autonomous driving
6 robots
7 Towards Artificial General Intelligence
Practice problems

References
Search

Detailed image
Detailed Image 1
GOODS SPECIFICS
- Date of issue: July 24, 2025
- Page count, weight, size: 536 pages | 1,059g | 188*257*21mm
- ISBN13: 9791173400070

You may also like

카테고리