Determines the optimal action for a policy (solved POMDP) for a given belief at a given epoch.
Arguments
- model
a solved POMDP.
- belief
The belief (probability distribution over the states) as a vector or a matrix with multiple belief states as rows. If
NULL, then the initial belief of the model is used.- epoch
what epoch of the policy should be used. Use 1 for converged policies.
Examples
data("Tiger")
Tiger
#> POMDP, list - Tiger Problem
#> Discount factor: 0.75
#> Horizon: Inf epochs
#> Size: 2 states / 3 actions / 2 obs.
#> Start: uniform
#> Solved: FALSE
#>
#> List components: ‘name’, ‘discount’, ‘horizon’, ‘states’, ‘actions’,
#> ‘observations’, ‘transition_prob’, ‘observation_prob’, ‘reward’,
#> ‘start’, ‘terminal_values’, ‘info’
sol <- solve_POMDP(model = Tiger)
# these are the states
sol$states
#> [1] "tiger-left" "tiger-right"
# belief that tiger is to the left
optimal_action(sol, c(1, 0))
#> [1] open-right
#> Levels: listen open-left open-right
optimal_action(sol, "tiger-left")
#> [1] open-right
#> Levels: listen open-left open-right
# belief that tiger is to the right
optimal_action(sol, c(0, 1))
#> [1] open-left
#> Levels: listen open-left open-right
optimal_action(sol, "tiger-right")
#> [1] open-left
#> Levels: listen open-left open-right
# belief is 50/50
optimal_action(sol, c(.5, .5))
#> [1] listen
#> Levels: listen open-left open-right
optimal_action(sol, "uniform")
#> [1] listen
#> Levels: listen open-left open-right
# the POMDP is converged, so all epoch give the same result.
optimal_action(sol, "tiger-right", epoch = 10)
#> [1] open-left
#> Levels: listen open-left open-right