Skip to contents

MDP Models

Define finite-state MDPs and inspect their states, actions, transitions, and rewards.

MDP() S() A() is_solved_MDP() is_converged_MDP() P_() R_()
Define an MDP Problem with Model Access
MDPSample()
Define an MDP With Only Sample Access
available_actions()
Available Actions in a State
transition_matrix() reward_matrix() start_vector() normalize_MDP()
Access to Parts of the Model Description
act()
Perform an Action
sample_MDP()
Sample Trajectories from an MDP
absorbing_states()
Absorbing States
normalize_state() normalize_state_id() normalize_state_label() normalize_state_features() normalize_action() normalize_action_id() normalize_action_label() state2features() features2state() s() get_state_features()
Conversions for Action and State IDs and Labels
find_reachable_states()
Find Reachable State Space from a Transition Model Function
reachable_states()
Find Reachable States
sample_MDP(<MDPSample>)
Sample Trajectories from an MDPSample
start(<MDPModel>) start(<MDPSample>)
Sample a Start State
transition_graph() plot_transition_graph() curve_multiple_directed()
Transition Graph
unreachable_states() remove_unreachable_states()
Unreachable States

Solvers

Solve MDPs with dynamic programming, linear programming, sampling, and reinforcement learning methods.

solve_MDP()
Solve an MDP Problem
solve_MDP_DP()
Solve MDPs using Dynamic Programming
solve_MDP_TD()
Solve MDPs using Tabular Temporal Differencing
solve_MDP_MC()
Solve MDPs using Monte Carlo Control
solve_MDP_LP()
Solve MDPs using Linear Programming
solve_MDP_SAMP()
Solve MDPs using Random-Sampling
solve_MDP_APPROX() approx_Q_value() approx_greedy_action() approx_greedy_policy() approx_V_plot()
Solve MDPs with Temporal Differencing with Function Approximation
solve_MDP_PG()
Solve MDPs with Policy Gradient Methods
schedule_exp() schedule_exp2() schedule_log() schedule_linear() schedule_harmonic()
Schedules to Reduce Alpha, Epsilon and Other Parameters
convergence_horizon()
Estimate the Convergence Horizon for an Infinite-Horizon MDP

Value Functions

Work with value functions.

Policies

Create and evaluate policies, choose actions, and calculate value functions and returns.

policy() add_policy() random_policy() manual_policy() induced_transition_matrix() induced_reward_matrix()
Extract, Create Add a Policy to a Model
action()
Choose an Action Given a Policy
expected_return()
Calculate the Expected Return of a Policy
greedy_action() greedy_policy()
Greedy Actions and Policies
policy_evaluation() policy_evaluation_LP() policy_evaluation_MC() policy_evaluation_bellman()
Policy Evaluation
regret() action_discrepancy() value_error()
Regret of a Policy and Related Measures
visit_probability()
State Visit Probability

Function Approximation

Approximate value functions and policies with state features and basis transformations.

Gridworlds

Create gridworld MDPs and inspect their layouts, transitions, and solutions.

Example MDPs

Explore the included maze, cliff-walking, and windy-gridworld examples.

Maze maze
Steward Russell's 4x3 Maze Gridworld MDP
Cliff_walking cliff_walking
Cliff Walking Gridworld MDP
Windy_gridworld windy_gridworld
Windy Gridworld MDP Windy Gridworld MDP
DynaMaze dynamaze
The Dyna Maze

Visualization

Plot gridworlds and transition graphs using the package’s color palettes.

Utilities

Helper functions.

colors_discrete() colors_continuous()
Default Colors for Visualization
round_stochastic()
Round a stochastic vector or a row-stochastic matrix