{"product_id":"138156","title":"Robust reinforcement learning ","description":"\u003ccenter\u003e\u003cdiv style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/tmgdisk01.cafe24.com\/images\/vs\/4172\/sv\/3jXPCfJfoTlRuKMPh0i7e3e10Veuil.png?v=1765060793\" style=\"max-width:100%;max-height:10px\"\u003e\u003c\/div\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\u003ccenter\u003e\n\n\u003cdiv style=\"width:95%\"\u003e\n\n\u003cdiv style=\"text-align:center;font-size:30px;font-weight:bolder;line-height:1.6em\"\u003e Robust reinforcement learning \u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"border-bottom:1px;border-bottom-style:dotted;border-color:;padding-bottom:20px\"\u003e\u003ccenter\u003e\u003ctable align=\"center\" width=\"100%\"\u003e\u003ctbody style=\"border:0px\"\u003e\n\n\u003ctr\u003e\u003ctd align=\"center\" style=\"line-height:1.2em;text-align:center;font-size:18px;color:black;font-weight:bold;padding-bottom:20px;\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\u003ctr\u003e\u003ctd style=\"text-align:center\"\u003e\u003cimg src=\"https:\/\/image.yes24.com\/goods\/89605439\/XL\" style=\"max-width:100%;height:auto\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\n\n\n\u003c\/tbody\u003e\u003c\/table\u003e\u003c\/center\u003e\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;{split_style6}padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e Description \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;word-break:break-all;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eBook Introduction\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\u003cdiv\u003e \u003cb\u003eThe absolute bible of reinforcement learning, revised for the first time in 20 years with significantly enhanced content!\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Reinforcement learning, one of the most actively researched areas in artificial intelligence, is a numerical computational learning method that maximizes the reward given to a learner interacting with a complex and uncertain environment.\u003cbr\u003e In their book, Robust Reinforcement Learning, Richard Sutton and Andrew Barto explain the core concepts and algorithms of reinforcement learning clearly and easily.\u003cbr\u003e Since the publication of the first edition, new topics have been added, and topics already covered have been updated with the latest content.\u003cbr\u003e\n\u003c\/div\u003e\u003c\/div\u003e\n\u003cdiv\u003e\u003cul\u003e  \u003cli\u003eYou can preview some of the book's contents.\u003cbr\u003e \u003cspan\u003ePreview\u003c\/span\u003e\n\n\u003c\/li\u003e\n\u003c\/ul\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eindex\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e \u003cb\u003eCHAPTER 01 Introduction 1\u003c\/b\u003e\u003cbr\u003e 1.1 Reinforcement Learning 2\u003cbr\u003e 1.2 Example 5\u003cbr\u003e 1.3 Components of Reinforcement Learning 7\u003cbr\u003e 1.4 Limitations and Scope 9\u003cbr\u003e 1.5 Extended Example: Tic-Tac-Toe 10\u003cbr\u003e 1.6 Summary 16\u003cbr\u003e 1.7 Early History of Reinforcement Learning 17\u003cbr\u003e Reference 27\u003cbr\u003e\u003cbr\u003e \u003cb\u003ePART I Solutions in Table Form\u003c\/b\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 02 Multiple Choice 31\u003c\/b\u003e\u003cbr\u003e 2.1 Multiple Choice Problem 32\u003cbr\u003e 2.2 Behavioral Value Method 34\u003cbr\u003e 2.3 10-choice test 35\u003cbr\u003e 2.4 Incremental Implementation 38\u003cbr\u003e 2.5 Traces of Abnormal Problems 40\u003cbr\u003e 2.6 Positive initial value 42\u003cbr\u003e 2.7 Trust Limit Action Selection 44\u003cbr\u003e 2.8 Gradient Multiple Selection Algorithm 46\u003cbr\u003e 2.9 Related Search (Contextual Multiple Choice) 50\u003cbr\u003e 2.10 Summary 51\u003cbr\u003e References and Historical Facts 54\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 03 Finite Markov Decision Processes 57\u003c\/b\u003e\u003cbr\u003e 3.1 Agent-Environment Interface 58\u003cbr\u003e 3.2 Goals and Rewards 64\u003cbr\u003e 3.3 Rewards and Episode 66\u003cbr\u003e 3.4 Unified Notation for Episodic and Continuous Work 69\u003cbr\u003e 3.5 Policy and Value Functions 70 \u003cbr\u003e3.6 Optimal Policy and Optimal Value Function 76\u003cbr\u003e 3.7 Optimality and Approximation 82\u003cbr\u003e 3.8 Summary 83\u003cbr\u003e References and Historical Facts 84\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 04 DYNAMIC PROGRAMMING 89\u003c\/b\u003e\u003cbr\u003e 4.1 Policy Evaluation (Prediction) 90\u003cbr\u003e 4.2 Policy Improvement 94\u003cbr\u003e 4.3 Policy Repetition 97\u003cbr\u003e 4.4 Value Repeat 100\u003cbr\u003e 4.5 Asynchronous Dynamic Programming 103\u003cbr\u003e 4.6 Generalized Policy Iteration 104\u003cbr\u003e 4.7 Efficiency of Dynamic Programming 106\u003cbr\u003e 4.8 Summary 107\u003cbr\u003e References and Historical Facts 109\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 05 Monte Carlo Methods 111\u003c\/b\u003e\u003cbr\u003e 5.1 Monte Carlo Forecasting 112\u003cbr\u003e 5.2 Monte Carlo Action Value Estimation 118\u003cbr\u003e 5.3 Monte Carlo Control 119\u003cbr\u003e 5.4 Monte Carlo Control Without Starting Exploration 123\u003cbr\u003e 5.5 Predicting Inactive Policies Using Importance Extraction 126\u003cbr\u003e 5.6 Incremental Implementation 133\u003cbr\u003e 5.7 Inactive Monte Carlo Control 135\u003cbr\u003e Importance Extraction Method Considering Discounts 138\u003cbr\u003e 5.9 Decision-Level Importance Extraction Method 139\u003cbr\u003e 5.10 Summary 141\u003cbr\u003e References and Historical Facts 143\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 06 Time-Differential Learning 145\u003c\/b\u003e\u003cbr\u003e 6.1 TD Prediction 146\u003cbr\u003e 6.2 Advantages of the TD Prediction Method 150\u003cbr\u003e 6.3 Optimality of TD(0) 153 \u003cbr\u003e6.4 Salsa: Active Policy TD Control 157\u003cbr\u003e 6.5 Q Learning: Inactive Policy TD Control 160\u003cbr\u003e 6.6 Expected value Salsa 162\u003cbr\u003e 6.7 Maximizing Variance and Dual Learning 163\u003cbr\u003e 6.8 Games, Post-Game Conditions, and Other Special Cases 166\u003cbr\u003e 6.9 Summary 168\u003cbr\u003e References and Historical Facts 169\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 07 n-step bootstrap 171\u003c\/b\u003e\u003cbr\u003e 7.1 n-step TD prediction 172\u003cbr\u003e 7.2 n-level salsa 177\u003cbr\u003e 7.3 Learning n-step inactive policies 179\u003cbr\u003e 7.4 Step-by-step method for decision making with control variables 181\u003cbr\u003e 7.5 Passive Policy Learning Without Importance Extraction: An n-Stage Tree Augmentation Algorithm 184\u003cbr\u003e 7.6 Integration Algorithm: n-Step Q(σ) 187\u003cbr\u003e 7.7 Summary 189\u003cbr\u003e References and Historical Facts 190\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 08 Planning and Learning Using Table-Based Methods 191\u003c\/b\u003e\u003cbr\u003e 8.1 Models and Plans 192\u003cbr\u003e 8.2 Dyna: Integrating Planning, Action, and Learning 194\u003cbr\u003e 8.3 When the model is wrong 199\u003cbr\u003e 8.4 Prioritized Batch Processing 202\u003cbr\u003e 8.5 Expected Updates vs. Sample Updates 206\u003cbr\u003e 8.6 Trajectory Sampling 210\u003cbr\u003e 8.7 Real-Time Dynamic Programming 213\u003cbr\u003e 8.8 Planning at the Decision Point 217 \u003cbr\u003e8.9 Empirical Exploration 219\u003cbr\u003e 8.10 Dice Rolling Algorithm 221\u003cbr\u003e 8.11 Monte Carlo Tree Search 223\u003cbr\u003e 8.12 Summary 227\u003cbr\u003e 8.13 Part 1 Summary: Dimension 228\u003cbr\u003e References and Historical Facts 231\u003cbr\u003e\u003cbr\u003e \u003cb\u003ePART II Approximate Solutions\u003c\/b\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 09 Active Policy Prediction Using Approximation 237\u003c\/b\u003e\u003cbr\u003e 9.1 Value Function Approximation 238\u003cbr\u003e 9.2 Predictive Objectives (VE) 239\u003cbr\u003e 9.3 Probabilistic Gradient and Semi-Gradient Methods 241\u003cbr\u003e 9.4 Linear Methods 246\u003cbr\u003e 9.5 Creating Features for Linear Methods 253\u003cbr\u003e 9.6 Manually selecting time interval parameters 268\u003cbr\u003e 9.7 Nonlinear Function Approximation: Artificial Neural Networks 269\u003cbr\u003e 9.8 Least Squares TD 275\u003cbr\u003e 9.9 Memory-Based Function Approximation 278\u003cbr\u003e 9.10 Kernel-Based Function Approximation 280\u003cbr\u003e A Deeper Look at 9\/11 Active Policy Learning: Focus and Emphasis 282\u003cbr\u003e 9.12 Summary 285\u003cbr\u003e References and Historical Facts 286\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 10 Active Policy Control with Approximation 293\u003c\/b\u003e\u003cbr\u003e 10.1 Episodic Semi-Slope Control 294\u003cbr\u003e 10.2 Semi-slope n-level Salsa 297\u003cbr\u003e 10.3 Average Reward: Setting a New Problem for Continuous Tasks 300 \u003cbr\u003e10.4 Objection to Discounted Settings 304\u003cbr\u003e 10.5 Differential Semi-Slope n-Step Salsa 307\u003cbr\u003e 10.6 Summary 308\u003cbr\u003e References and Historical Facts 308\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 11: Inactive Policy Methods Using Approximation 311\u003c\/b\u003e\u003cbr\u003e 11.1 Semi-slope Method 312\u003cbr\u003e 11.2 Example of Inactive Policy Emission 315\u003cbr\u003e 11.3 Deadly Trinity 320\u003cbr\u003e 11.4 Linear Value Function Geometry 322\u003cbr\u003e 11.5 Gradient Descent in Bellman Error 327\u003cbr\u003e 11.6 Bellman error cannot be learned 332\u003cbr\u003e 11.7 Gradient TD Method 337\u003cbr\u003e 11.8 Strong TD Method 341\u003cbr\u003e 11.9 Reducing Variance 343\u003cbr\u003e 11.10 Summary 345\u003cbr\u003e References and Historical Facts 346\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 12 Eligible Traces 349\u003c\/b\u003e\u003cbr\u003e 12.1 λ gain 350\u003cbr\u003e 12.2 TD(λ) 355\u003cbr\u003e 12.3 Interrupted n-step λ gain method 359\u003cbr\u003e 12.4 Re-Updating: Online λ Gain Algorithm 361\u003cbr\u003e 12.5 True Online TD(λ) 363\u003cbr\u003e 12.6 Dutch Traces in Monte Carlo Learning 366\u003cbr\u003e 12.7 Salsa (λ) 368\u003cbr\u003e 12.8 Variable λ and γ 372\u003cbr\u003e 12.9 Inactive Policy Traces with Control Variables 374\u003cbr\u003e 12.10 From Watkins' Q(λ) to Tree Augmentation(λ) 378 \u003cbr\u003e12.11 Stable Inactivity Policy Method Using Traces 381\u003cbr\u003e 12.12 Implementation Issue 383\u003cbr\u003e 12.13 Conclusion 384\u003cbr\u003e References and Historical Facts 386\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 13 Policy Gradient Methods 389\u003c\/b\u003e\u003cbr\u003e 13.1 Policy Approximation and Its Advantages 390\u003cbr\u003e 13.2 Policy Gradient Summary 393\u003cbr\u003e 13.3 REINFORCE: Monte Carlo Policy Gradient 395\u003cbr\u003e 13.4 REINFORCE 399 with reference values\u003cbr\u003e 13.5 The Doer-Critic Method 401\u003cbr\u003e 13.6 Policy Gradients for Continuous Problems 403\u003cbr\u003e 13.7 Policy Parameterization for Continuous Actions 406\u003cbr\u003e 13.8 Summary 408\u003cbr\u003e References and Historical Facts 409\u003cbr\u003e\u003cbr\u003e \u003cb\u003ePART III A DEEPER DIG\u003c\/b\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 14 Psychology 413\u003c\/b\u003e\u003cbr\u003e 14.1 Prediction and Control 414\u003cbr\u003e 14.2 Classical Conditioning 416\u003cbr\u003e 14.3 Instrumental Conditioning 433\u003cbr\u003e 14.4 Delayed Reinforcement 438\u003cbr\u003e 14.5 Cognitive Map 440\u003cbr\u003e 14.6 Habitual and Goal-Directed Behavior 442\u003cbr\u003e 14.7 Summary 447\u003cbr\u003e References and Historical Facts 449\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 15: NEUROSCIENCE 457\u003c\/b\u003e\u003cbr\u003e 15.1 Neuroscience Fundamentals 458\u003cbr\u003e 15.2 Reward Signals, Reinforcement Signals, Value, and Prediction Error 460\u003cbr\u003e 15.3 Reward Prediction Error Hypothesis 463 \u003cbr\u003e15.4 Dopamine 465\u003cbr\u003e 15.5 Experimental Support for the Reward Prediction Error Hypothesis 469\u003cbr\u003e 15.6 TD error\/dopamine similarity 473\u003cbr\u003e 15.7 Neurobehavior-Critic 479\u003cbr\u003e 15.8 Doer and Critic Learning Rules 482\u003cbr\u003e 15.9 Hedonistic Neurons 488\u003cbr\u003e 15.10 Collective Reinforcement Learning 490\u003cbr\u003e 15.11 Model-Based Methods in the Brain 494\u003cbr\u003e 15.12 Addiction 496\u003cbr\u003e 15.13 Summary 497\u003cbr\u003e References and Historical Facts 501\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 16 Applications and Case Studies 511\u003c\/b\u003e\u003cbr\u003e 16.1 TD-Garmon 511\u003cbr\u003e 16.2 Samuel's Checkers Player 518\u003cbr\u003e 16.3 Watson's Double Wager 522\u003cbr\u003e 16.4 Memory Control Optimization 526\u003cbr\u003e 16.5 Human-level video game skills 531\u003cbr\u003e 16.6 Mastering the Game of Baduk 539\u003cbr\u003e 16.7 Personalized Web Services 550\u003cbr\u003e 16.8 Fever Rise 554\u003cbr\u003e\u003cbr\u003e \u003cb\u003eCHAPTER 17 Frontier 559\u003c\/b\u003e\u003cbr\u003e 17.1 General Value Functions and Auxiliary Operations 559\u003cbr\u003e 17.2 Temporal Abstraction through Options 562\u003cbr\u003e 17.3 Observations and Conditions 565\u003cbr\u003e 17.4 Design of the Compensation Signal 572\u003cbr\u003e 17.5 Remaining Issues 576\u003cbr\u003e 17.6 The Future of Artificial Intelligence 580\u003cbr\u003e References and Historical Facts 584\u003cbr\u003e\u003cbr\u003e Reference 588\u003cbr\u003e Search 626\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e  \u003ch5\u003e\n\u003cb\u003eDetailed image\u003c\/b\u003e \u003c\/h5\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e\u003cdiv\u003e\u003cimg src=\"https:\/\/image.yes24.com\/momo\/TopCate2997\/MidCate009\/299683116(1).jpg\" border=\"0\" alt=\"Detailed Image 1\"\u003e\u003c\/div\u003e\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003eInto the book\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e Artificial intelligence technology has advanced tremendously in the 20 years since this book was first published in 1998.\u003cbr\u003e Advances in machine learning technologies, including reinforcement learning, have provided a major impetus for the development of artificial intelligence.\u003cbr\u003e While the advancement of machine learning technology has been partly due to the remarkable advancement of computers' computing power, the development of new theories and algorithms has also played a significant role.\u003cbr\u003e Despite these changes, work on the second edition of this book was delayed for a long time, and work could not begin until 2012.\u003cbr\u003e The purpose of this second edition is no different from that of the first publication of this book.\u003cbr\u003e That is, the goal is to enable readers from all relevant fields to easily and clearly understand the core concepts and algorithms of reinforcement learning.\u003cbr\u003e --- From the \"Preface\"\u003cbr\u003e\u003cbr\u003e Consider the following learning problem: \u003cbr\u003eYou must repeatedly choose one of k different options or actions.\u003cbr\u003e After each selection, a numerical reward is given.\u003cbr\u003e At this time, the value representing the reward is obtained from a stationary probability distribution (a probability distribution that does not change over time_translator) determined according to the selected action.\u003cbr\u003e The goal of selection is to maximize the expected value of the total amount of reward given over a period of time, for example, over the period of selecting actions 1,000 times or over 1,000 time steps.\u003cbr\u003e --- p.32\u003cbr\u003e\u003cbr\u003e Another reasonable answer is to simply observe that we have encountered state A once and the resulting payoff was zero, so we estimate the value of V(A) to be zero.\u003cbr\u003e This answer is given by the batch Monte Carlo method.\u003cbr\u003e Note that this is the answer that derives the least squares error for the training data. \u003cbr\u003eIn fact, this answer yields an error of 0 on the training data.\u003cbr\u003e --- p.155\u003cbr\u003e\u003cbr\u003e Overfitting is a problem in all function approximation methods that have many degrees of freedom and fit a function based on limited training data.\u003cbr\u003e Although this problem is less pronounced in online reinforcement learning, which is not constrained by limited training data, effective generalization remains a critical issue.\u003cbr\u003e Overfitting is a problem with ANNs in general, but it becomes more serious with deep ANNs because of their tendency to have a very large number of weights.\u003cbr\u003e --- p.272\u003cbr\u003e\u003cbr\u003e In contrast to trial-phase models such as the Rescorla-Wagner model, the TD model is a real-time model. \u003cbr\u003eIn the Rescorla-Wagner model, a single step t represents an entire conditioning trial. The TD model does not care about the details of what happens during the time a conditioning trial occurs or what happens between conditioning trials.\u003cbr\u003e During each conditioning trial, the animal may experience different stimuli occurring at specific times and for specific periods of time.\u003cbr\u003e\n\n\u003c\/div\u003e\n\u003cdiv\u003e --- p.423\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cbr\u003e\u003cdiv\u003e\u003ch5\u003e \u003cb\u003ePublisher's Review\u003c\/b\u003e\n\u003c\/h5\u003e\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e \u003cb\u003eThe absolute bible of reinforcement learning, revised for the first time in 20 years with significantly enhanced content!\u003cbr\u003e Understand the core concepts and latest algorithms of reinforcement learning in a simple and clear way!\u003c\/b\u003e\u003cbr\u003e\u003cbr\u003e Reinforcement learning, one of the most actively researched areas in artificial intelligence, is a numerical computational learning method that maximizes the reward given to a learner interacting with a complex and uncertain environment. \u003cbr\u003eIn their book, Robust Reinforcement Learning, Richard Sutton and Andrew Barto explain the core concepts and algorithms of reinforcement learning clearly and easily.\u003cbr\u003e Since the publication of the first edition, new topics have been added, and topics already covered have been updated with the latest content.\u003cbr\u003e\u003cbr\u003e Like the first edition, the second edition focuses on core online learning algorithms, but includes more mathematical content in separate text boxes.\u003cbr\u003e This book is broadly divided into three parts:\u003cbr\u003e\u003cbr\u003e ■ In the first part, we covered as many reinforcement learning methods as possible, applying only table-based methods that can find accurate solutions.\u003cbr\u003e Many of the algorithms presented in the first part are new to the second edition, including UCB, Expected Value Salsa, and Dual Learning. \u003cbr\u003e■ In the second part, the methods presented in the first part are extended to function approximation-based methods, with new sections covering topics such as artificial neural networks and Fourier-based methods, and the content on inactive policy learning and policy gradient methods is enriched.\u003cbr\u003e ■ The third part includes new chapters on how reinforcement learning relates to psychology and neuroscience, and updated chapters on case studies such as AlphaGo and AlphaGo Zero, Atari games, and IBM Watson's betting strategies.\u003cbr\u003e In the final chapter, we discussed the impact of reinforcement learning on future society. \u003cbr\u003e\n\n\u003c\/div\u003e\n\u003cdiv\u003e\u003c\/div\u003e\n\u003c\/div\u003e\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cdiv style=\"width:95%;padding-top:20px;padding-bottom:20px\"\u003e\n\n\u003cdiv style=\"text-align:left;font-size:16px;font-weight:bold;padding-bottom:20px\"\u003e GOODS SPECIFICS \u003c\/div\u003e\n\n\u003cdiv style=\"text-align:left;font-size:14px;line-height:1.6em;\"\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eDate of issue:\u003c\/strong\u003e March 31, 2020\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003ePage count, weight, size:\u003c\/strong\u003e 664 pages | 1,290g | 188*245*33mm\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN13:\u003c\/strong\u003e 9791190665179\u003c\/div\u003e\n\n\u003cdiv style=\"width:100%;margin-bottom:5px;line-height:1.6em;font-size:14px\"\u003e - \u003cstrong\u003eISBN10:\u003c\/strong\u003e 1190665174 \u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\n\u003c\/div\u003e\n\n\u003ccenter\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003ccenter\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"height:10px\"\u003e\u003c\/td\u003e\u003c\/tr\u003e\u003c\/table\u003e\u003c\/center\u003e\n\n\u003cspan\u003e\u003c\/span\u003e\n\n\u003c\/center\u003e\n\n\n\u003c\/center\u003e","brand":"LIBRAIRIE COREENNE","offers":[{"title":"Default Title","offer_id":43893206057002,"sku":"138156","price":51.0,"currency_code":"EUR","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0683\/2750\/5962\/files\/7f0db25de1bd521544a455703f76a69d.jpg?v=1765391912","url":"https:\/\/librairie.coreenne.fr\/en\/products\/138156","provider":"LIBRAIRIE COREENNE","version":"1.0","type":"link"}