Calculates the state visit probability (the modified stationary distribution) when following a policy from the start state.
Arguments
- model
a solved MDP object.
- pi
the used policy. If missing the policy in
modelis used.- start
specification of the start distribution. If missing the specification in
modelis used.- method
calculate the modified stationary distribution using
"power"(power iteration) or"sample"(trajectory sampling).- ...
further arguments are passed on to the internal implementations (see Details section).
Details
The visit probability is the stationary distribution for the transition matrix induced by the policy. To account for absorbing states, we modify the transition matrix by setting all outgoing probabilities from absorbing states to 0.
Power iteration
The stationary distribution can be estimated as the sum of multiplying the start distribution repeatedly with the modified transition matrix induced by the policy. We stop multiplying when the largest difference between entries in the two consecutive vectors is less then the extra parameter:
min_errstop criterion for the power iteration (default:1e-6).sparselogical; should a sparse transition matrix be used for the power iteration?
The resulting vector is normalized to probabilities.
See also
Other policy:
action(),
expected_return(),
greedy_action(),
policy(),
policy_evaluation(),
regret()
Examples
data("Maze")
Maze
#> MDPModel, MDP - Stuart Russell's 3x4 Maze
#> Discount factor: 1
#> Horizon: Inf epochs
#> Size: 4 actions / 11 states
#> Storage: transition prob as matrix / reward as matrix. Total size: 28.8 Kb
#> Start: s(3,1)
#> Model list components: ‘name’, ‘discount’, ‘horizon’, ‘states’,
#> ‘actions’, ‘start’, ‘transition_model’, ‘reward’, ‘info’,
#> ‘absorbing_states’
sol <- solve_MDP(Maze)
visit_probability(sol)
#> s(1,1) s(2,1) s(3,1) s(1,2) s(3,2) s(1,3)
#> 0.162710379 0.183049179 0.162710382 0.162710373 0.020338798 0.160481419
#> s(2,3) s(3,3) s(1,4) s(2,4) s(3,4)
#> 0.017831260 0.000000000 0.128385085 0.001783124 0.000000000
# gw_matrix also can calculate the visit_probability.
gw_matrix(sol, what = "visit_probability")
#> [,1] [,2] [,3] [,4]
#> [1,] 0.1627104 0.1627104 0.16048142 0.128385085
#> [2,] 0.1830492 NA 0.01783126 0.001783124
#> [3,] 0.1627104 0.0203388 0.00000000 0.000000000