Reinforcement Learning. An Introduction [2nd ed.]
Book information
Description
Contents......Page 3 Preface......Page 9 Notation......Page 15 Reinforcement Learning......Page 19 Examples......Page 22 Elements of Reinforcement Learning......Page 24 Limitations and Scope......Page 25 An Extended Example: Tic-Tac-Toe......Page 26 Early History of Reinforcement Learning......Page 31 --- Tabular Solution Methods......Page 41 A k-armed Bandit Problem......Page 42 Action-value Methods......Page 44 The 10-armed Testbed......Page 45 Incremental Implementation......Page 47 Tracking a Nonstationary Problem......Page 49 Optimistic Initial Values......Page 51 Upper-Confidence-Bound Action Selection......Page 52 Gradient Bandit Algorithms......Page 54 Associative Search (Contextual Bandits)......Page 58 Summary......Page 59 The Agent–Environment Interface......Page 63 Goals and Rewards......Page 69 Returns and Episodes......Page 70 Unified Notation for Episodic and Continuing Tasks......Page 73 Policies and Value Functions......Page 74 Optimal Policies and Optimal Value Functions......Page 78 Optimality and Approximation......Page 83 Summary......Page 84 Dynamic Programming......Page 88 Policy Evaluation (Prediction)......Page 89 Policy Improvement......Page 91 Policy Iteration......Page 95 Value Iteration......Page 97 Asynchronous Dynamic Programming......Page 100 Generalized Policy Iteration......Page 101 Efficiency of Dynamic Programming......Page 102 Summary......Page 103 Monte Carlo Methods......Page 106 Monte Carlo Prediction......Page 107 Monte Carlo Estimation of Action Values......Page 111 Monte Carlo Control......Page 112 Monte Carlo Control without Exploring Starts......Page 115 Off-policy Prediction via Importance Sampling......Page 118 Incremental Implementation......Page 124 Off-policy Monte Carlo Control......Page 125 *Discounting-aware Importance Sampling......Page 127 *Per-decision Importance Sampling......Page 129 Summary......Page 130 TD Prediction......Page 133 Advantages of TD Prediction Methods......Page 138 Optimality of TD(0)......Page 140 Sarsa: On-policy TD Control......Page 143 Q-learning: Off-policy TD Control......Page 145 Expected Sarsa......Page 147 Maximization Bias and Double Learning......Page 148 Games, Afterstates, and Other Special Cases......Page 150 Summary......Page 152 n-Step Bootstrapping......Page 155 n-step TD Prediction......Page 156 n-step Sarsa......Page 159 n-step Off-policy Learning......Page 162 *Per-decision Methods with Control Variates......Page 164 Off-policy Learning Without Importance Sampling: The n-step Tree Backup Algorithm......Page 166 *A Unifying Algorithm: n-step Q()......Page 168 Summary......Page 171 Models and Planning......Page 173 Dyna: Integrated Planning, Acting, and Learning......Page 175 When the Model Is Wrong......Page 180 Prioritized Sweeping......Page 182 Expected vs. Sample Updates......Page 186 Trajectory Sampling......Page 188 Real-time Dynamic Programming......Page 191 Planning at Decision Time......Page 194 Heuristic Search......Page 195 Rollout Algorithms......Page 197 Monte Carlo Tree Search......Page 199 Summary of the Chapter......Page 202 Summary of Part I: Dimensions......Page 203 --- Approximate Solution Methods......Page 208 On-policy Prediction with Approximation......Page 210 Value-function Approximation......Page 211 The Prediction Objective (VE)......Page 212 Stochastic-gradient and Semi-gradient Methods......Page 213 Linear Methods......Page 217 Polynomials......Page 223 Fourier Basis......Page 224 Coarse Coding......Page 228 Tile Coding......Page 230 Radial Basis Functions......Page 234 Selecting Step-Size Parameters Manually......Page 235 [23pt][l]9.7Nonlinear Function Approximation: Artificial Neural Networks......Page 236 Least-Squares TD......Page 241 Memory-based Function Approximation......Page 243 Kernel-based Function Approximation......Page 245 Looking Deeper at On-policy Learning: Interest and Emphasis......Page 247 Summary......Page 249 Episodic Semi-gradient Control......Page 255 Semi-gradient n-step Sarsa......Page 259 Average Reward: A New Problem Setting for Continuing Tasks......Page 261 Deprecating the Discounted Setting......Page 265 Differential Semi-gradient n-step Sarsa......Page 267 Summary......Page 268 Off-policy Methods with Approximation......Page 269 Semi-gradient Methods......Page 270 Examples of Off-policy Divergence......Page 272 The Deadly Triad......Page 276 Linear Value-function Geometry......Page 278 Gradient Descent in the Bellman Error......Page 281 The Bellman Error is Not Learnable......Page 286 Gradient-TD Methods......Page 290 Emphatic-TD Methods......Page 293 Reducing Variance......Page 295 Summary......Page 296 Eligibility Traces......Page 299 The -return......Page 300 TD()......Page 304 n-step Truncated -return Methods......Page 307 Redoing Updates: Online -return Algorithm......Page 309 True Online TD()......Page 311 *Dutch Traces in Monte Carlo Learning......Page 313 Sarsa()......Page 315 Variable and......Page 319 *Off-policy Traces with Control Variates......Page 321 Watkins's Q() to Tree-Backup()......Page 324 Stable Off-policy Methods with Traces......Page 326 Implementation Issues......Page 328 Conclusions......Page 329 Policy Gradient Methods......Page 333 Policy Approximation and its Advantages......Page 334 The Policy Gradient Theorem......Page 336 REINFORCE: Monte Carlo Policy Gradient......Page 338 REINFORCE with Baseline......Page 341 Actor–Critic Methods......Page 343 Policy Gradient for Continuing Problems......Page 345 Policy Parameterization for Continuous Actions......Page 347 Summary......Page 349 --- Looking Deeper......Page 351 Psychology......Page 352 Prediction and Control......Page 353 Classical Conditioning......Page 354 Blocking and Higher-order Conditioning......Page 356 The Rescorla–Wagner Model......Page 357 The TD Model......Page 360 TD Model Simulations......Page 361 Instrumental Conditioning......Page 368 Delayed Reinforcement......Page 372 Cognitive Maps......Page 374 Habitual and Goal-directed Behavior......Page 375 Summary......Page 379 Neuroscience......Page 388 Neuroscience Basics......Page 389 Reward Signals, Reinforcement Signals, Values, and Prediction Errors......Page 391 The Reward Prediction Error Hypothesis......Page 392 Dopamine......Page 394 [23pt][l]15.5Experimental Support for the Reward Prediction Error Hypothesis......Page 398 TD Error/Dopamine Correspondence......Page 401 Neural Actor–Critic......Page 406 Actor and Critic Learning Rules......Page 409 Hedonistic Neurons......Page 413 Collective Reinforcement Learning......Page 415 Model-based Methods in the Brain......Page 418 Addiction......Page 420 Summary......Page 421 TD-Gammon......Page 431 Samuel's Checkers Player......Page 436 Watson's Daily-Double Wagering......Page 439 Optimizing Memory Control......Page 442 Human-level Video Game Play......Page 446 Mastering the Game of Go......Page 451 AlphaGo......Page 454 AlphaGo Zero......Page 457 Personalized Web Services......Page 460 Thermal Soaring......Page 463 General Value Functions and Auxiliary Tasks......Page 468 Temporal Abstraction via Options......Page 470 Observations and State......Page 473 Designing Reward Signals......Page 478 Remaining Issues......Page 481 The Future of Artificial Intelligence......Page 484 Refs......Page 490 Index......Page 528
Similar books
Reinforcement Learning, second edition: An Introduction (Solutions) (Instructor's Solution Manual)
2018 · PDF
Reinforcement Learning: An Introduction
2018 · PDF
Reinforcement Learning: An Introduction
1998 · PDF
Reinforcement Learning. An Introduction
2002 · PDF
Reinforcement Learning: An Introduction
1998 · PDF
Reinforcement Learning: An Introduction
1998 · PDF
Reinforcement learning
1998 · CHM
Reinforcement Learning - An Introduction
1998 · PDF