{"product_id":"139618","title":"Robust deep reinforcement learning ","description":"\u003ccenter\u003e\u003cdiv style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/tmgdisk01.cafe24.com\/images\/vs\/4172\/sv\/3jXPBvi5kFeTRzIDrtDeqF9Nz6DPhX.png?v=1765075579\" style=\"max-width:100%;max-height:10px\"\u003e\u003c\/div\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\n\n\u003cdiv style=\"width:95%\"\u003e\n\n\u003cdiv style=\"text-align:center;font-size:30px;font-weight:bolder;line-height:1.6em\"\u003e Robust deep reinforcement learning \u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"border-bottom:1px;border-bottom-style:dotted;border-color:;padding-bottom:20px\"\u003e\u003ccenter\u003e\u003ctable align=\"center\" width=\"100%\"\u003e\u003ctbody style=\"border:0px\"\u003e\n\n\u003ctr\u003e\u003ctd align=\"center\" style=\"line-height:1.2em;text-align:center;font-size:18px;color:black;font-weight:bold;padding-bottom:20px;\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\u003ctr\u003e\u003ctd style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/image.yes24.com\/goods\/106709909\/XL\" style=\"max-width:100%;height:auto\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\n\u003c\/tbody\u003e\u003c\/table\u003e\u003c\/center\u003e\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;{split_style6}padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e Description \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;word-break:break-all;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eBook Introduction\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\u003cdiv\u003e \u003cb\u003eThe perfect way to build a solid foundation in deep reinforcement learning!\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e This book is an introduction to deep reinforcement learning that uniquely combines theory and practice.\u003cbr\u003e It begins with an intuitive explanation, moves on to a detailed explanation of deep reinforcement learning algorithms and implementation methods using the SLM Lab library, and finally covers the details for applying deep reinforcement learning in practice.\u003cbr\u003e\n\n\u003c\/div\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cul\u003e\u003cli\u003e You can preview some of the book's contents.\u003cbr\u003e \u003cspan\u003ePreview\u003c\/span\u003e\n\n\u003c\/li\u003e\u003c\/ul\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eindex\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e  \u003cdiv\u003eTranslator's Preface xii\u003cbr\u003e Beta Reader Review xiii\u003cbr\u003e Recommendation xv\u003cbr\u003e Beginning xvi\u003cbr\u003e Acknowledgements xxi\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 01 Introduction to Reinforcement Learning 1\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 1.1 Reinforcement Learning 1\u003cbr\u003e 1.2 Reinforcement Learning as an MDP 7\u003cbr\u003e 1.3 Functions Learned in Reinforcement Learning 11\u003cbr\u003e 1.4 Deep Reinforcement Learning Algorithm 13\u003cbr\u003e 1.4.1 Policy-Based Algorithm 14\u003cbr\u003e 1.4.2 Value-Based Algorithms 15\u003cbr\u003e 1.4.3 Model-Based Algorithms 16\u003cbr\u003e 1.4.4 Combined Method 17\u003cbr\u003e 1.4.5 Algorithms covered in this book 18\u003cbr\u003e 1.4.6 Active and Inactive Policy Algorithms 19\u003cbr\u003e 1.4.7 Summary 19\u003cbr\u003e 1.5 Deep Learning for Reinforcement Learning 20\u003cbr\u003e 1.6 Reinforcement Learning and Supervised Learning 22\u003cbr\u003e 1.6.1 Absence of Oracle 23\u003cbr\u003e 1.6.2 The Scarcity of Feedback 24\u003cbr\u003e 1.6.3 Data Generation 24\u003cbr\u003e 1.7 Summary 25\u003cbr\u003e\u003cbr\u003e \u003cb\u003ePART I Policy-Based and Value-Based Algorithms\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 02 REINFORCE 29\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 2.1 Policy 30\u003cbr\u003e 2.2 Objective Function 31\u003cbr\u003e 2.3 Policy Slope 31\u003cbr\u003e 2.3.1 Policy Gradient Calculation 33\u003cbr\u003e 2.4 Monte Carlo Sampling 36\u003cbr\u003e 2.5 REINFORCE Algorithm 37\u003cbr\u003e 2.5.1 Improved REINFORCE 38\u003cbr\u003e 2.6 REINFORCE Implementation 39 \u003cbr\u003e2.6.1 Minimal REINFORCE Implementation 39\u003cbr\u003e 2.6.2 Creating Policies with PyTorch 42\u003cbr\u003e 2.6.3 Action Extraction 44\u003cbr\u003e 2.6.4 Policy Loss Calculation 45\u003cbr\u003e 2.6.5 REINFORCE Training Loop 46\u003cbr\u003e 2.6.6 Active Policy Replay Memory 47\u003cbr\u003e 2.7 Training of REINFORCE Agents 50\u003cbr\u003e 2.8 Experimental Results 53\u003cbr\u003e 2.8.1 Experiment: The Effect of Discount Rate ?? 53\u003cbr\u003e 2.8.2 Experiment: The Effect of Reference Values ​​55\u003cbr\u003e 2.9 Summary 57\u003cbr\u003e 2.10 Further Reading 57\u003cbr\u003e 2.11 History 58\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 03 SARSA 59\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 3.1 Q function and V function 60\u003cbr\u003e 3.2 Time-lapse learning 63\u003cbr\u003e 3.2.1 Intuition about temporal difference learning 66\u003cbr\u003e 3.3 Salsa's Action Selection 73\u003cbr\u003e 3.3.1 Exploration and Utilization 74\u003cbr\u003e 3.4 Salsa Algorithm 75\u003cbr\u003e 3.4.1 Activation Policy Algorithm 76\u003cbr\u003e 3.5 Application of Salsa 77\u003cbr\u003e 3.5.1 Behavior Function: Epsilon-Greedy 77\u003cbr\u003e 3.5.2 Calculating Q Loss 78\u003cbr\u003e 3.5.3 Salsa Training Loop 80\u003cbr\u003e 3.5.4 Active Policy Deployment Reproduction Memory 81\u003cbr\u003e 3.6 Salsa Agent Training 83\u003cbr\u003e 3.7 Experimental Results 86\u003cbr\u003e 3.7.1 Experiment: The Effect of Learning Rate 86\u003cbr\u003e 3.8 Summary 87\u003cbr\u003e 3.9 Further Reading 88\u003cbr\u003e 3.10 History 89\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 04 Deep Q Network (DQN) 91\u003c\/b\u003e\u003cbr\u003e \u003cbr\u003e4.1 Learning the Q Function of DQN 92\u003cbr\u003e 4.2 Action Selection in DQN 94\u003cbr\u003e 4.2.1 Boltzmann Policy 97\u003cbr\u003e 4.3 Experience Reproduction 100\u003cbr\u003e 4.4 DQN Algorithm 101\u003cbr\u003e 4.5 Application of DQN 103\u003cbr\u003e 4.5.1 Calculating Q Loss 103\u003cbr\u003e 4.5.2 DQN Training Loop 104\u003cbr\u003e 4.5.3 Reproduced Memory 105\u003cbr\u003e 4.6 Training the DQN Agent 108\u003cbr\u003e 4.7 Experimental Results 111\u003cbr\u003e 4.7.1 Experiment: The Effect of Neural Network Architecture 111\u003cbr\u003e 4.8 Summary 113\u003cbr\u003e 4.9 Further Reading 114\u003cbr\u003e 4.10 History 114\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 05 Improved DQN 115\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 5.1 Target Network 116\u003cbr\u003e 5.2 Dual DQN 119\u003cbr\u003e 5.3 Prioritized Experience Reproduction (PER) 123\u003cbr\u003e 5.3.1 Importance Sampling 125\u003cbr\u003e 5.4 Implementation of the Modified DQN 126\u003cbr\u003e 5.4.1 Network Initialization 127\u003cbr\u003e 5.4.2 Calculating Q Loss 128\u003cbr\u003e 5.4.3 Target Network Update 129\u003cbr\u003e 5.4.4 DQN 130 with target network\u003cbr\u003e 5.4.5 Dual DQN 130\u003cbr\u003e 5.4.6 Reproducing Prioritized Experiences 131\u003cbr\u003e 5.5 Training a DQN Agent for Atari Games 137\u003cbr\u003e 5.6 Experimental Results 142\u003cbr\u003e 5.6.1 Experiment: The Effect of Dual DQN and PER 142\u003cbr\u003e 5.7 Summary 146\u003cbr\u003e 5.8 Further Reading 146\u003cbr\u003e\u003cbr\u003e \u003cb\u003ePART II Combined Methods\u003c\/b\u003e\u003cbr\u003e \u003cbr\u003e\u003cb\u003eCHAPTER 06 Advantage Doer-Critic (A2C) 149\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 6.1 Actor 150\u003cbr\u003e 6.2 Critics 150\u003cbr\u003e 6.2.1 Advantage Function 151\u003cbr\u003e 6.2.2 Learning about the Advantage Function 155\u003cbr\u003e 6.3 A2C Algorithm 156\u003cbr\u003e 6.4 Implementation of A2C 159\u003cbr\u003e 6.4.1 Advantage Estimation 160\u003cbr\u003e 6.4.2 Calculating Value Loss and Policy Loss 162\u003cbr\u003e 6.4.3 Doer-Critic Training Loop 163\u003cbr\u003e 6.5 Network Architecture 164\u003cbr\u003e 6.6 Training the A2C Agent 166\u003cbr\u003e 6.6.1 Applying A2C with n-step gain to the Pong game 166\u003cbr\u003e 6.6.2 Applying A2C to Pong Game Using GAE 169\u003cbr\u003e 6.6.3 A2C 170 using n-step gain in the bipedal pedestrian problem\u003cbr\u003e 6.7 Experimental Results 173\u003cbr\u003e 6.7.1 Experiment: The Effect of n-Stage Gains 173\u003cbr\u003e 6.7.2 Experiment: The Effect of ?? on GAE 175\u003cbr\u003e 6.8 Summary 176\u003cbr\u003e 6.9 Further Reading 177\u003cbr\u003e 6.10 History 177\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 07 Proximal Policy Optimization (PPO) 179\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 7.1 Proxy Purpose 180\u003cbr\u003e 7.1.1 Performance Collapse 180\u003cbr\u003e 7.1.2 Modifying the Objective Function 182\u003cbr\u003e 7.2 Proximal Policy Optimization (PPO) 189\u003cbr\u003e 7.3 PPO Algorithm 193\u003cbr\u003e 7.4 Implementation of PPO 195\u003cbr\u003e 7.4.1 Calculating PPO Policy Losses 195 \u003cbr\u003e7.4.2 PPO Training Loop 196\u003cbr\u003e 7.5 Training of PPO Agents 198\u003cbr\u003e 7.5.1 PPO 198 for Pong Game\u003cbr\u003e 7.5.2 PPO 201 for two-legged pedestrians\u003cbr\u003e 7.6 Experimental Results 203\u003cbr\u003e 7.6.1 Experiment: The Effect of ?? on GAE 204\u003cbr\u003e 7.6.2 Experiment: Effect of the Clipping Variable ?? 205\u003cbr\u003e 7.7 Summary 207\u003cbr\u003e 7.8 Further Reading 208\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER PARALLELIZATION METHODS 209\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 8.1 Synchronous Parallelism 210\u003cbr\u003e 8.2 Asynchronous Parallelism 212\u003cbr\u003e 8.2.1 Hogwild! 213\u003cbr\u003e 8.3 Training the A3C Agent 216\u003cbr\u003e 8.4 Summary 219\u003cbr\u003e 8.5 Further Reading 219\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 09 Algorithm Summary 221\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e \u003cb\u003ePART III Details for Practice\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 10 Working with Deep Reinforcement Learning 225\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 10.1 Software Engineering Techniques 226\u003cbr\u003e 10.1.1 Unit Tests 226\u003cbr\u003e 10.1.2 Code Quality 232\u003cbr\u003e 10.1.3 Git Workflow 233\u003cbr\u003e 10.2 Debugging Tips 236\u003cbr\u003e 10.2.1 Survival Signal 236\u003cbr\u003e 10.2.2 Diagnosis of Policy Slope 237\u003cbr\u003e 10.2.3 Diagnosis of Data 238\u003cbr\u003e 10.2.4 Preprocessor 239\u003cbr\u003e 10.2.5 Memory 239\u003cbr\u003e 10.2.6 Algorithm Function 240\u003cbr\u003e 10.2.7 Neural Networks 240\u003cbr\u003e 10.2.8 Algorithm Simplification 243 \u003cbr\u003e10.2.9 Simplifying the Problem 243\u003cbr\u003e 10.2.10 Hyperparameter 244\u003cbr\u003e 10.2.11 Lab Workflow 244\u003cbr\u003e 10.3 Atari Trick 245\u003cbr\u003e 10.4 Deep Reinforcement Learning Almanac 249\u003cbr\u003e 10.4.1 Hyperparameter Table 249\u003cbr\u003e 10.4.2 Algorithm Performance Comparison 252\u003cbr\u003e 10.5 Summary 255\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 11 SLM Lab 257\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 11.1 Algorithms Implemented in SLM Lab 257\u003cbr\u003e 11.2 spec file 260\u003cbr\u003e 11.2.1 Search Specification Syntax 262\u003cbr\u003e 11.3 Running SLM Lab 265\u003cbr\u003e 11.3.1 SLM Lab Command 265\u003cbr\u003e 11.4 Analysis of Experimental Results 266\u003cbr\u003e 11.4.1 Overview of Experimental Data 266\u003cbr\u003e 11.5 Summary 268\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 12 NETWORK ARCHITECTURE 269\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 12.1 Types of Neural Networks 269\u003cbr\u003e 12.1.1 Multilayer Perceptron (MLP) 270\u003cbr\u003e 12.1.2 Convolutional Neural Networks (CNNs) 272\u003cbr\u003e 12.1.3 Recurrent Neural Networks (RNNs) 274\u003cbr\u003e 12.2 Guide to Selecting Network Groups 275\u003cbr\u003e 12.2.1 MDP and POMDP 275\u003cbr\u003e 12.2.2 Selecting a Network for Your Environment 279\u003cbr\u003e 12.3 Net API 282\u003cbr\u003e 12.3.1 Estimation of Input and Output Layer Shapes 284\u003cbr\u003e 12.3.2 Automatic Network Creation 286\u003cbr\u003e 12.3.3 Training Step 289\u003cbr\u003e 12.3.4 Exposure of base methods 290\u003cbr\u003e 12.4 Summary 291 \u003cbr\u003e12.5 Further Reading 292\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 13 HARDWARE 293\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 13.1 Computer 294\u003cbr\u003e 13.2 Data Types 300\u003cbr\u003e 13.3 Data Type Optimization in Reinforcement Learning 302\u003cbr\u003e 13.4 Hardware Selection 307\u003cbr\u003e 13.5 Summary 308\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 14 STATUS 311\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Example 312 of State 14.1\u003cbr\u003e 14.2 Completeness of State 319\u003cbr\u003e 14.3 State Complexity 320\u003cbr\u003e 14.4 Loss of State Information 325\u003cbr\u003e 14.4.1 Image Grayscaling 325\u003cbr\u003e 14.4.2 Dioxide 326\u003cbr\u003e 14.4.3 Hash Dispatch 327\u003cbr\u003e 14.4.4 Metadata Loss 327\u003cbr\u003e 14.5 Preprocessing 331\u003cbr\u003e 14.5.1 Standardization 332\u003cbr\u003e 14.5.2 Image Processing 333\u003cbr\u003e 14.5.3 Temporal Preprocessing 335\u003cbr\u003e 14.6 Summary 339\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 15: BEHAVIOR 341\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 15.1 Example of Action 341\u003cbr\u003e 15.2 Completeness of Action 345\u003cbr\u003e 15.3 The Complexity of Behavior 347\u003cbr\u003e 15.4 Summary 352\u003cbr\u003e 15.5 Further Reading: Designing Behaviors in Everyday Life 353\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 16 REWARDS 357\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 16.1 The Role of Compensation 357\u003cbr\u003e 16.2 Guidelines for Compensation Design 359\u003cbr\u003e 16.3 Summary 364\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 17 Transfer Functions 365\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e 17.1 Feasibility Check 366\u003cbr\u003e 17.2 Reality Check 368\u003cbr\u003e 17.3 Summary 371\u003cbr\u003e\u003cbr\u003e APPENDIX A: Deep Reinforcement Learning Timeline 372 \u003cbr\u003eAPPENDIX B Environment Example 374\u003cbr\u003e B.1 Discrete Environments 375\u003cbr\u003e B.1.1 CartPole-v0 375\u003cbr\u003e B.1.2 MountainCar-v0 376\u003cbr\u003e B.1.3 LunarLander-v2 377\u003cbr\u003e B.1.4 PongNoFrameskip-v4 378\u003cbr\u003e B.1.5 BreakoutNoFrameskip-v4 378\u003cbr\u003e B.2 Continuous Environment 379\u003cbr\u003e B.2.1 Pendulum-v0 379\u003cbr\u003e B.2.2 BipedalWalker-v2 380\u003cbr\u003e\u003cbr\u003e Epilogue 381\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eDetailed image\u003c\/b\u003e \u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cdiv\u003e\u003cimg src=\"https:\/\/image.yes24.com\/momo\/TopCate3750\/MidCate002\/374915889(1).jpg\" border=\"0\" alt=\"Detailed Image 1\"\u003e\u003c\/div\u003e\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eInto the book\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e This book introduces the entire process of deep reinforcement learning.\u003cbr\u003e It begins by presenting intuition, then explains the theory and algorithms, and concludes with actual implementations and practical advice.\u003cbr\u003e This is precisely why this book includes a library called SLM Lab, which contains implementation code for all the algorithms covered in this book.\u003cbr\u003e In short, this book is exactly what we wish we had when we first started studying deep reinforcement learning.\u003cbr\u003e\u003cbr\u003e --- p.18\u003cbr\u003e \u003cbr\u003eSince 2012, deep learning has been successfully applied to a variety of problems and has contributed to the development of cutting-edge technologies in a wide range of fields, including computer vision, machine translation, natural language understanding, and speech synthesis.\u003cbr\u003e As I write this, deep learning is the most powerful function approximation technique humans have ever created.\u003cbr\u003e\u003cbr\u003e --- p.20\u003cbr\u003e\u003cbr\u003e Deep reinforcement learning algorithms typically have many hyperparameters.\u003cbr\u003e For example, the type of network, architecture, activation function, optimization technique, and learning rate must be determined.\u003cbr\u003e More advanced neural network functions may include gradient clipping and learning rate decay schemes, which are only relevant in the 'deep' part of deep reinforcement learning!\u003cbr\u003e --- p.50\u003cbr\u003e\u003cbr\u003e The name 'Monte Carlo' doesn't mean anything in particular. \u003cbr\u003eIt is simply remembered as another expression for 'probabilistic estimation'.\u003cbr\u003e But the origin of the name is interesting.\u003cbr\u003e The name was suggested by Nicholas Metropolis, a physicist and computer designer who also coined the odd name MANIAC computer.\u003cbr\u003e Metropolis heard the story of Ulam's uncle who borrowed money from his relatives simply because he 'had to go to Monte Carlo'.\u003cbr\u003e After that, the name Monte Carlo seemed too appropriate to denote probabilistic estimation.\u003cbr\u003e --- p.58\u003cbr\u003e\u003cbr\u003e When designing a new reinforcement learning algorithm or component, it is necessary to prove that the designed algorithm is theoretically correct before implementation.\u003cbr\u003e This is especially true when conducting research. \u003cbr\u003eSimilarly, when trying to solve a problem in a new environment, it is necessary to first confirm whether the problem is truly solvable with reinforcement learning before applying the algorithm.\u003cbr\u003e This should be taken into consideration especially when developing applications.\u003cbr\u003e If a reinforcement learning algorithm fails even though everything is theoretically correct and can be solved with reinforcement learning, it may be due to an implementation error.\u003cbr\u003e If so, you need to debug your code.\u003cbr\u003e\u003cbr\u003e --- p.226\u003cbr\u003e\u003cbr\u003e In practice, forming higher-level patterns is equivalent to creating a control strategy using the original controls.\u003cbr\u003e This is similar to how chords become simpler control strategies constructed from individual piano keys.\u003cbr\u003e This technique is a kind of meta control, and people can do it at any time. \u003cbr\u003eCurrently, there is no way to enable reinforcement learning agents to design their own control strategies.\u003cbr\u003e Therefore, meta control is required for the agent.\u003cbr\u003e From the agent's perspective, we need to design higher-level patterns from a human perspective.\u003cbr\u003e Sometimes, instead of expressing an action as a single, complex action, it can be expressed more simply by combining several sub-actions.\u003cbr\u003e\n\n\u003c\/div\u003e\n\u003cdiv\u003e --- p.348 \u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e GOODS SPECIFICS \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003ePublication date:\u003c\/strong\u003e February 17, 2022\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003ePage count, weight, size:\u003c\/strong\u003e 428 pages | 188*245*21mm\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN13:\u003c\/strong\u003e 9791191600674\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN10:\u003c\/strong\u003e 119160067X \u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cspan\u003e\u003c\/span\u003e\n\n\u003c\/center\u003e\n\n\n\u003c\/center\u003e","brand":"LIBRAIRIE COREENNE","offers":[{"title":"Default Title","offer_id":43893369503786,"sku":"139618","price":40.0,"currency_code":"EUR","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0683\/2750\/5962\/files\/602f8279bb0a31e33f374881fe829fed.jpg?v=1765398695","url":"https:\/\/librairie.coreenne.fr\/en\/products\/139618","provider":"LIBRAIRIE COREENNE","version":"1.0","type":"link"}