Skip to contents

Calculates the regret of a policy relative to a benchmark policy.

Usage

regret(policy, benchmark, start = NULL)

Arguments

policy

a solved POMDP containing the policy to calculate the regret for.

benchmark

a solved POMDP with the (optimal) policy. Regret is calculated relative to this policy.

start

the used start (belief) state. If NULL then the start (belief) state of the benchmark is used.

Value

the regret as a difference of expected long-term rewards.

Details

Regret is defined as \(V^{\pi^*}(s_0) - V^{\pi}(s_0)\) with \(V^\pi\) representing the expected long-term state value (represented by the value function) given the policy \(\pi\) and the start state \(s_0\). For POMDPs the start state is the start belief \(b_0\).

Note that for regret usually the optimal policy \(\pi^*\) is used as the benchmark. Since the optimal policy may not be known, regret relative to the best known policy can be used.

Author

Michael Hahsler

Examples

data(Tiger)

sol_optimal <- solve_POMDP(Tiger)
sol_optimal
#> POMDP, list - Tiger Problem
#>   Discount factor: 0.75
#>   Horizon: Inf epochs
#>   Size: 2 states / 3 actions / 2 obs.
#>   Start: uniform
#>   Solved:
#>     Method: ‘grid’
#>     Solution converged: TRUE
#>     # of alpha vectors: 5
#>     Total expected reward: 1.933439
#> 
#>   List components: ‘name’, ‘discount’, ‘horizon’, ‘states’, ‘actions’,
#>     ‘observations’, ‘transition_prob’, ‘observation_prob’, ‘reward’,
#>     ‘start’, ‘info’, ‘solution’

# perform exact value iteration for 10 epochs
sol_quick <- solve_POMDP(Tiger, method = "enum", horizon = 10)
sol_quick
#> POMDP, list - Tiger Problem
#>   Discount factor: 0.75
#>   Horizon: 10 epochs
#>   Size: 2 states / 3 actions / 2 obs.
#>   Start: uniform
#>   Solved:
#>     Method: ‘enum’
#>     Solution converged: FALSE
#>     # of alpha vectors: 160
#>     Total expected reward: 1.661560
#> 
#>   List components: ‘name’, ‘discount’, ‘horizon’, ‘states’, ‘actions’,
#>     ‘observations’, ‘transition_prob’, ‘observation_prob’, ‘reward’,
#>     ‘start’, ‘info’, ‘solution’

regret(sol_quick, benchmark = sol_optimal)
#> [1] 0.2718789