Exploration and exploitation with Markov decision processes

The aim of this project is to understand the interplay between randomness and optimization in the choice of a Markov decision process and of a policy, working with simple examples such as gridworlds. Particular attention will be paid to entropy (maximization and regularization).