Hi, I have gone through the explanation about Q learning (http://gradientdescending.com/q-learning-example-with-liars-dice-in-r/), but feel a little bit confused about the idea of probability bucket. What does it use for and how to understand it?
PS: When I try to run the simulation, there is always an error: object 'Q.mat' not found.
Thank you
Hi, I have gone through the explanation about Q learning (http://gradientdescending.com/q-learning-example-with-liars-dice-in-r/), but feel a little bit confused about the idea of probability bucket. What does it use for and how to understand it?
PS: When I try to run the simulation, there is always an error: object 'Q.mat' not found.
Thank you