Not implemented yet: Solve the MDP control problem using a parameterized policy and policy gradient methods. The implemented method one-step actor-critic control and actor-critic control with eligibility traces.
Usage
solve_MDP_PG(
model,
method = "actor-critic",
horizon = NULL,
discount = NULL,
alpha_actor = schedule_exp(0.2, 0.1),
alpha_critic = schedule_exp(0.2, 0.1),
epsilon = schedule_exp(1, 0.1),
lambda = 0,
n,
w = NULL,
theta = NULL,
...,
matrix = TRUE,
continue = FALSE,
progress = TRUE,
verbose = FALSE
)Arguments
- model
an MDP problem specification.
- method
string; one of the following solution methods:
'sarsa'- horizon
an integer with the number of epochs for problems with a finite planning horizon. If set to
Inf, the algorithm continues running iterations till it converges to the infinite horizon solution. IfNULL, then the horizon specified inmodelwill be used.- discount
discount factor in range \((0, 1]\). If
NULL, then the discount factor specified inmodelwill be used.- alpha_actor, alpha_critic
alpha schedules
- epsilon
used for the \(\epsilon\)-greedy behavior policies. A scalar value between 0 and 1 or a schedule.
- lambda
the trace-decay parameter for the an accumulating trace. If
lambda = 0then 1-step Sarsa is used.- n
number of episodes used for learning.
- w
an initial weight vector. By default a vector with 0s is used.
- theta
parameter...
- ...
further parameters are passed on to the solver function.
- matrix
logical; if
TRUEthen matrices for the transition model and the reward function are taken from the model first. This can be slow if functions need to be converted or do not fit into memory if the models are large. If these components are already matrices, then this is very fast. ForFALSE, the transition probabilities and the reward is extracted when needed. This is slower, but removes the time and memory requirements needed to calculate the matrices.- continue
logical; show a progress bar with estimated time for completion.
- progress
logical; show a progress bar with estimated time for completion.
- verbose
logical or a numeric verbose level; if set to
TRUEor1, the function displays the used algorithm parameters and progress information. Levels>1provide more detailed solver output in the R console.
Value
solve_MDP() returns an object of class MDP or MDPSample which is a list with the
model specifications (model), the solution (solution).
The solution is a list with the elements that depend on the used method. Common
elements are:
methodwith the name of the used methodparameters used.
convergeddid the algorithm converge (NA) for finite-horizon problems.policya list representing the policy graph. The list only has one element for converged solutions.
References
Sutton, Richard S., and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. Second. The MIT Press. http://incompleteideas.net/book/the-book-2nd.html.
See also
Other solver:
convergence_horizon(),
schedule,
solve_MDP(),
solve_MDP_APPROX(),
solve_MDP_DP(),
solve_MDP_LP(),
solve_MDP_MC(),
solve_MDP_SAMP(),
solve_MDP_TD()
Other MDPSample:
MDPSample(),
absorbing_states(),
act(),
action_state_helpers,
reachable_states(),
sample_MDP.MDPSample(),
start