Estimate Q-learning convergence time to epsilon-optimal policy given state-action count, discount factor, and learning rate schedule.
Enter the values for the patient or scenario you are assessing. Estimate Q-learning convergence time to epsilon-optimal policy given state-action count, discount factor, and learning rate schedule. Use the q-learning convergence result to inform your clinical assessment.