Vikram KrishnamurthyCambridge University Press, 3/21/2016EAN 9781107134607, ISBN10: 1107134609Hardcover, 488 pages, 25.4 x 18 x 2.5 cmLanguage: EnglishCovering formulation, algorithms, and structural results, and linking theory to real-world applications in controlled sensing (including social learning, adaptive radars and sequential detection), this book focuses on the conceptual foundations of partially observed Markov decision processes (POMDPs). It emphasizes structural results in stochastic dynamic programming, enabling graduate students and researchers in engineering, operations research, and economics to understand the underlying unifying themes without getting weighed down by mathematical technicalities. Bringing together research from across the literature, the book provides an introduction to nonlinear filtering followed by a systematic development of stochastic dynamic programming, lattice programming and reinforcement learning for POMDPs. Questions addressed in the book include: when does a POMDP have a threshold optimal policy? When are myopic policies optimal? How do local and global decision makers interact in adaptive decision making in multi-agent social learning where there is herding and data incest? And how can sophisticated radars and sensors adapt their sensing in real time?Preface1. IntroductionPart I. Stochastic Models and Bayesian Filtering2. Stochastic state-space models3. Optimal filtering4. Algorithms for maximum likelihood parameter estimation5. Multi-agent sensingsocial learning and data incestPart II. Partially Observed Markov Decision Processes. Models and Algorithms6. Fully observed Markov decision processes7. Partially observed Markov decision processes (POMDPs)8. POMDPs in controlled sensing and sensor schedulingPart III. Partially Observed Markov Decision Processes9. Structural results for Markov decision processes10. Structural results for optimal filters11. Monotonicity of value function for POMPDs12. Structural results for stopping time POMPDs13. Stopping time POMPDs for quickest change detection14. Myopic policy bounds for POMPDs and sensitivity to model parametersPart IV. Stochastic Approximation and Reinforcement Learning15. Stochastic optimization and gradient estimation16. Reinforcement learning17. Stochastic approximation algorithmsexamples18. Summary of algorithms for solving POMPDsAppendix A. Short primer on stochastic simulationAppendix B. Continuous-time HMM filtersAppendix C. Markov processesAppendix D. Some limit theoremsBibliographyIndex.