A structured online programme from IISc provide a strong foundation in Reinforcement Learning through the various tools, techniques and algorithms used as well as to cover the state-of-the-art algorithms in Deep Reinforcement Learning involving simulation-based neural network methods.
Professor in the Department of Computer Science and Automation, Convenor of Stochastic Systems Laboratory, Member, Steering Group, Robert Bosch Centre for Cyber Physical Systems, Associate Faculty, Center for Infrastructure, Sustainable Transportation and Urban Planning in Indian Institute of Science.His major research interests lie in the area of stochastic approximation, with emphasis on algorithms for control and optimization of stochastic dynamic systems, in particular reinforcement learning and simulation optimization. Amongst the application domains, I am interested in vehicular traffic control, autonomous systems, communication and wireless networks, and smart grids.
Introduces Reinforcement Learning (RL) as techniques combining optimal control, simulation/data-driven optimization, and approximation methods for dynamic decision-making under uncertainty.
Highlights applications in areas such as Adaptive Control, Signal Processing, Manufacturing, Communication and Wireless Networks, Autonomous Systems, and Data Mining.
Explains model-free algorithms, which learn without prior knowledge of system dynamics or protocols.
Provides a strong foundation in RL concepts, tools, techniques, and algorithms.
Covers state-of-the-art Deep Reinforcement Learning methods using simulation-based neural network approaches.
Introduction to Reinforcement Learning – examples and applications
Multi-armed Bandits – action selection strategies
Multi-armed Bandits – algorithms; Introduction to Markov Decision Processes
Markov Decision Processes – Examples, formulations
Numerical approaches for Markov Decision Processes
Monte-Carlo model-free Reinforcement Learning Algorithms for prediction
Monte-Carlo Algorithms for Control; Temporal Difference Methods
One and n-Step Temporal Difference Learning, Q-learning, SARSA, Expected SARSA, Double Q-learning
Function Approximation Methods, TD Learning/SARSA with Linear Function Approximation
Neural network architectures, Deep Q-learning
Introduction to policy gradient methods – basic principles and results
Policy gradient algorithms – REINFORCE, Actor-Critic
Apply online at iisc.online · New batches every semester (Jan–May and Aug–Dec)