back to search
You will learn the mathematical foundations of reinforcement learning: Markov decision processes (MDPs), tabular RL methods such as Monte Carlo, Temporal Difference, SARSA and Q-Learning, as well as basics of stochastic approximation to analyze the convergence of these algorithms. In the end you can formulate dynamic decision problems under uncertainty as MDPs and solve them with tabular RL algorithms.
No ratings for this module yet.
Only fill in the categories you can judge – for each one, either stars and text together or nothing at all.
Reviews are automatically checked before they are published.
Official page in TUMonline · Details are not binding.