Algorithm API
Dynamic programming
gym_classics2.algorithms.dynamic_programming
This file implements dynamic programming algorithms for solving Markov Decision Processes (MDPs) in gym-classics environments with model access. The algorithms include value iteration and policy iteration, which are fundamental methods in reinforcement learning for computing optimal policies and value functions.
backup
Computes the Bellman backup for a given state and action.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
A gym-classics environment with model access. |
required | |
discount
|
The discount factor. |
required | |
V
|
The current value function. |
required | |
state
|
The current state. |
required | |
action
|
The action to evaluate. |
required |
Returns: The computed Q-value for the given state and action.
Source code in gym_classics2/algorithms/dynamic_programming.py
value_iteration
Performs value iteration for the given environment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
A gym-classics environment with model access. |
required | |
discount
|
The discount factor (0 <= discount <= 1). |
required | |
precision
|
The precision for convergence (default: 1e-3). |
0.001
|
|
history
|
If True, returns a list of intermediate value functions. |
False
|
|
verbose
|
If True, prints progress information. |
False
|
Returns:
| Type | Description |
|---|---|
|
The optimal value function V. If history is True, returns a list of intermediate value functions. |
Source code in gym_classics2/algorithms/dynamic_programming.py
policy_evaluation
Evaluates a given policy to compute its value function.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
A gym-classics environment with model access. |
required | |
discount
|
The discount factor (0 <= discount <= 1). |
required | |
policy
|
The policy to evaluate. |
required | |
precision
|
The precision for convergence (default: 1e-3). |
0.001
|
|
max_backups
|
Maximum number of backups to perform to prevent infinite loops (default: 1000). |
1000
|
Returns:
| Type | Description |
|---|---|
|
The value function for the given policy. |
Source code in gym_classics2/algorithms/dynamic_programming.py
policy_improvement
Improves the policy based on the given value function.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
policy
|
The current policy to improve. |
required | |
V_policy
|
The value function of the current policy. |
required | |
env
|
A gym-classics environment with model access. |
required | |
discount
|
The discount factor (0 <= discount <= 1). |
required | |
precision
|
The precision for determining stability (default: 1e-3). |
0.001
|
Returns:
| Type | Description |
|---|---|
|
A tuple (improved_policy, stable) where stable is True if the policy did not change. |
Source code in gym_classics2/algorithms/dynamic_programming.py
policy_iteration
policy_iteration(env, discount, precision=0.001, max_backups=1000, history=False, verbose=False, rng=None)
Performs policy iteration for the given environment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
A gym-classics environment with model access. |
required | |
discount
|
The discount factor (0 <= discount <= 1). |
required | |
precision
|
The precision for convergence (default: 1e-3). |
0.001
|
|
max_backups
|
Maximum number of iterations used in policy evaluation. Note: this prevents an infinite loop for policies that do not reach a terminal state. |
1000
|
|
history
|
If True, returns lists of intermediate policies and value functions. |
False
|
|
verbose
|
If True, prints progress information. |
False
|
|
rng
|
NumPy generator or integer seed used to initialize the policy. |
None
|
Returns:
| Type | Description |
|---|---|
|
The optimal policy. If history is True, returns a tuple (policy_list, V_list) containing lists of intermediate policies and value functions. |
Source code in gym_classics2/algorithms/dynamic_programming.py
Policy helpers
gym_classics2.algorithms.policy
This file implements different policy representations and functions for working with policies in gym-classics environments. It includes functions for creating random policies, encoding policies for display, and computing greedy policies based
make_multidiscrete_policy
Converts a tabular policy vector to a multi-discrete tabular policy stored in a dictionary that can be used for sample.
Source code in gym_classics2/algorithms/policy.py
random_policy
Create a random policy for the given environment.
The policy is represented as a numpy array where each entry corresponds to an action for a state.
rng may be a NumPy generator or an integer seed.
Source code in gym_classics2/algorithms/policy.py
encode_policy
Encode a policy for display. The policy is represented as a numpy array where each entry corresponds to an action for a state. The function returns a list of action names corresponding to the actions in the policy.
Source code in gym_classics2/algorithms/policy.py
greedy_policy
Calculate a greedy policy from a state-value function.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
V
|
One-dimensional state-value array. |
required | |
env
|
Environment with a discrete state space and model access. |
required | |
discount
|
Discount factor in |
1
|
|
rng
|
NumPy generator or integer seed for random tie-breaking. |
None
|
Returns:
| Type | Description |
|---|---|
|
Integer action ID selected for each state. |
Source code in gym_classics2/algorithms/policy.py
greedy_policy_Q
Calculate a greedy policy from an action-value function.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
Q
|
Two-dimensional action-value array indexed by state and action. |
required | |
env
|
Environment with discrete observation and action spaces. |
required | |
discount
|
Unused; retained for API compatibility with |
1
|
|
rng
|
NumPy generator or integer seed for random tie-breaking. |
None
|
Returns:
| Type | Description |
|---|---|
|
Integer action ID selected for each state. |
Source code in gym_classics2/algorithms/policy.py
epsilon_greedy_action
Select an epsilon-greedy action from tabular action values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
policy
|
Action values for all actions, optionally indexed first by state. |
required | |
state
|
Current state index. If omitted, |
None
|
|
epsilon
|
Probability of selecting a uniformly random action. |
0
|
|
rng
|
NumPy generator or integer seed. |
None
|
Returns:
| Type | Description |
|---|---|
|
Selected integer action ID. |
Source code in gym_classics2/algorithms/policy.py
Monte Carlo methods
gym_classics2.algorithms.monte_carlo_methods
This file implements tabular Monte Carlo methods for policy evaluation and control in gym-classics environments with discrete state spaces.
sample_episode
sample_episode(env, policy=None, start_state=None, start_action=None, epsilon=0, max_len=1000, verbose=False, rng=None)
Samples an episode from the environment using the given policy and starting conditions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
The environment to sample from. |
required | |
policy
|
A mapping from states to actions. For discrete state spaces, use
an action list in state order. If |
None
|
|
start_state
|
The state to start the episode from (if None, the environment's default starting state will be used). |
None
|
|
start_action
|
The action to take in the first step of the episode (if None, the action will be chosen according to the policy or randomly if no policy is given). |
None
|
|
epsilon
|
The probability of taking a random action instead of the policy's action at each step (for epsilon-greedy exploration). |
0
|
|
max_len
|
The maximum length of the episode to prevent infinite loops. |
1000
|
|
verbose
|
If True, prints the state transitions and rewards for each step in the episode. |
False
|
|
rng
|
NumPy generator or integer seed for policy and exploration choices. |
None
|
Returns:
| Type | Description |
|---|---|
|
A list of (state, action, reward, next_state) tuples representing the episode. |
|
|
The episode ends when a terminal state is reached or when max_len steps have been taken. |
|
|
Note that the last tuple in the episode will have a next_state that is either terminal or the |
|
|
state at which the episode was truncated due to max_len. |
Source code in gym_classics2/algorithms/monte_carlo_methods.py
on_policy_state_distribution
Estimate a policy's state distribution using rng for exploration.
Source code in gym_classics2/algorithms/monte_carlo_methods.py
MC_prediction
Estimate a policy's state values with first-visit Monte Carlo prediction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
policy
|
Array-like mapping from state IDs to action IDs. |
required | |
env
|
Gymnasium environment with a discrete observation space. |
required | |
discount
|
Reward discount factor. |
required | |
n
|
Number of episodes to sample. |
100
|
|
max_episode_len
|
Maximum sampled steps per episode. |
100
|
|
verbose
|
Print episode progress when true. |
False
|
|
rng
|
NumPy generator or integer seed for episode sampling. |
None
|
Returns:
| Type | Description |
|---|---|
|
A NumPy value array with one entry per state. Unvisited states contain |
|
|
|
Source code in gym_classics2/algorithms/monte_carlo_methods.py
MC_control_ES_textbook
MC_control_ES_textbook(env, discount, n=100, Q=None, max_episode_len=100, history=False, verbose=False, rng=None)
Monte Carlo control with exploring starts and stored sample returns.
This direct textbook implementation stores every first-visit return. Prefer
:func:MC_control_ES for larger experiments because its running-average
update uses less memory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
A tabular |
required | |
discount
|
Reward discount factor. |
required | |
n
|
Number of episodes to sample. |
100
|
|
Q
|
Optional initial action-value array. |
None
|
|
max_episode_len
|
Maximum sampled steps per episode. |
100
|
|
history
|
Retain intermediate policies, Q arrays, episodes, and returns. |
False
|
|
verbose
|
Print episode details; values greater than one print transitions. |
False
|
|
rng
|
NumPy generator or integer seed for all algorithm choices. |
None
|
Returns:
| Type | Description |
|---|---|
|
|
Source code in gym_classics2/algorithms/monte_carlo_methods.py
163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 | |
MC_control_ES
MC_control_ES(env, discount, n=100, Q=None, max_episode_len=100, history=False, verbose=False, rng=None)
Monte Carlo Control with Exploring Starts (incremental version). This algorithm estimates the optimal action-value function Q and the corresponding greedy policy by sampling episodes with exploring starts. It uses incremental updates to compute the average returns for each (s,a) pair, which is more memory efficient than storing all returns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
A tabular |
required | |
discount
|
Reward discount factor. |
required | |
n
|
Number of episodes to sample. |
100
|
|
Q
|
Optional initial action-value array. |
None
|
|
max_episode_len
|
Maximum sampled steps per episode. |
100
|
|
history
|
Retain intermediate policies, Q arrays, episodes, and returns. Note: retaining the history may require a lot of memory. |
False
|
|
verbose
|
Print episode details; values greater than one print transitions. |
False
|
|
rng
|
NumPy generator or integer seed for all algorithm choices. |
None
|
Returns:
| Type | Description |
|---|---|
|
|
Source code in gym_classics2/algorithms/monte_carlo_methods.py
249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 | |
Temporal-difference learning
gym_classics2.algorithms.temporal_difference_learning
This file implements temporal difference learning algorithms for policy evaluation and control in gym-classics
environments with discrete state spaces. The algorithms include Sarsa(0) and Q-learning, which are fundamental
methods in reinforcement learning for learning value functions and optimal policies from experience without requiring a
model of the environment.
Sarsa_0
Learn action values with one-step on-policy Sarsa.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Gymnasium environment with discrete observation and action spaces. |
required | |
discount
|
Reward discount factor in |
required | |
alpha
|
Scalar step size or |
required | |
epsilon
|
Scalar exploration probability or |
required | |
Q
|
Optional initial action-value array shaped
|
None
|
|
n
|
Number of training episodes. |
100
|
|
verbose
|
Print individual updates when true. |
False
|
|
history
|
Retain Q arrays, discounted episode returns, and episode lengths. |
False
|
|
rng
|
NumPy generator or integer seed for exploration and tie-breaking. |
None
|
Returns:
| Type | Description |
|---|---|
|
The learned Q array. If |
|
|
the dictionary contains |
Raises:
| Type | Description |
|---|---|
AssertionError
|
If the observation space is not discrete. |
Source code in gym_classics2/algorithms/temporal_difference_learning.py
Q_learning
Learn action values with one-step off-policy Q-learning.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Gymnasium environment with discrete observation and action spaces. |
required | |
discount
|
Reward discount factor in |
required | |
alpha
|
Scalar step size or |
required | |
epsilon
|
Scalar exploration probability or |
required | |
Q
|
Optional initial action-value array shaped
|
None
|
|
n
|
Number of training episodes. |
100
|
|
verbose
|
Print progress information when true. |
False
|
|
history
|
Retain Q arrays, discounted episode returns, episode lengths, and state-visit counts. |
False
|
|
rng
|
NumPy generator or integer seed for exploration and tie-breaking. |
None
|
Returns:
| Type | Description |
|---|---|
|
The learned Q array. If |
|
|
the dictionary contains |
|
|
|
Raises:
| Type | Description |
|---|---|
AssertionError
|
If the observation space is not discrete. |
Source code in gym_classics2/algorithms/temporal_difference_learning.py
105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 | |
Linear function approximation
gym_classics2.algorithms.linear_approximation
Linear function approximation algorithms for policy evaluation and control.
The algorithms do not require discrete state spaces. Callers provide a
state_features(state, env) function that converts states to state feature vectors.
state_action_features
Construct the state-action feature vector x(s, a).
The state feature vector is expected to have the form
[1, x1, ..., xd], with a leading intercept. The returned vector keeps
one shared intercept and has a separate block of the remaining state
features for each action. Only the intercept and the weights for a are
selected; all other elements are set to zero.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
s
|
State to represent. |
required | |
a
|
Integer ID of the action to represent. |
required | |
env
|
Environment providing the discrete action space. |
required | |
state_features
|
Callable that returns |
required |
Returns:
| Type | Description |
|---|---|
|
Block-coded NumPy feature vector for the state-action pair |
Source code in gym_classics2/algorithms/linear_approximation.py
v_hat
Compute the linear state-value approximation v_hat(s, w).
Approximates the state value as w^T x(s), the weighted sum of the components in the
state feature vector x(s). The weight and feature vectors must have the
same length.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
s
|
State to evaluate. |
required | |
w
|
Weight vector of the linear approximator. |
required | |
env
|
Environment containing the state. |
required | |
state_features
|
Callable returning the feature vector |
required |
Returns:
| Type | Description |
|---|---|
|
Scalar estimate of the expected return from |
Source code in gym_classics2/algorithms/linear_approximation.py
q_hat
Compute the linear action-value approximation q_hat(s, a, w).
Estimates the q-value as w^T x(s, a) using the block-coded state-action
features produced by state_action_features(). The weight vector must
have the same length as that feature vector.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
s
|
State to evaluate. |
required | |
a
|
Integer ID of the action to evaluate. |
required | |
w
|
Weight vector of the linear approximator. |
required | |
env
|
Environment providing the discrete action space. |
required | |
state_features
|
Callable returning a state feature vector with a
leading intercept for |
required |
Returns:
| Type | Description |
|---|---|
|
Scalar estimate of the expected return from taking |
Source code in gym_classics2/algorithms/linear_approximation.py
epsilon_greedy_action_w
Select an epsilon-greedy action from approximate action values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
s
|
Current state. |
required | |
w
|
Weight vector for the action-value approximator. |
required | |
env
|
Environment providing the discrete action space. |
required | |
state_features
|
Callable converting |
required | |
epsilon
|
Probability of selecting a uniformly random action. |
0
|
|
rng
|
NumPy generator or integer seed. |
None
|
Returns:
| Type | Description |
|---|---|
|
Selected integer action ID. |
Source code in gym_classics2/algorithms/linear_approximation.py
MSVE
Calculate the weighted mean squared value error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
V
|
Estimated value for each state. |
required | |
V_true
|
Reference value for each state. |
required | |
weight
|
Weight for each state, typically its stationary visitation probability. If omitted, use unit weights. |
None
|
Returns:
| Type | Description |
|---|---|
|
Weighted sum of squared value errors. |
Source code in gym_classics2/algorithms/linear_approximation.py
semi_gradient_TD0_estimation
semi_gradient_TD0_estimation(env, state_features, policy, n, alpha, gamma, max_episode_length=1000, verbose=False)
Estimate state values with semi-gradient TD(0).
This function runs TD(0) learning with function approximation over multiple episodes generated from a given policy and environment. Updates are performed using the semi-gradient of the value function approximation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Episodic Gymnasium environment used to generate experience. |
required | |
state_features
|
Callable converting |
required | |
policy
|
Deterministic policy indexed by state. |
required | |
n
|
Number of training episodes. |
required | |
alpha
|
Step size or schedule. |
required | |
gamma
|
Discount factor in |
required | |
max_episode_length
|
Maximum number of steps per episode. |
1000
|
|
verbose
|
Whether to print step-by-step diagnostics. |
False
|
Returns:
| Type | Description |
|---|---|
|
Learned weight vector for the approximate value function. |
Source code in gym_classics2/algorithms/linear_approximation.py
semi_gradient_Sarsa_0
semi_gradient_Sarsa_0(env, state_features, n, epsilon, alpha, gamma, w=None, max_episode_length=1000, verbose=False, history=False, rng=None)
Run semi-gradient Sarsa(0) with function approximation.
Implements the semi-gradient Sarsa(0) algorithm for estimating the optimal action-value function q_*(s, a) using a differentiable function approximator q̂(s, a, w). Actions are selected according to an ε-greedy policy derived from the current action-value estimate.
Episodes are truncated after max_episode_length time steps.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Episodic environment used to generate experience. |
required | |
state_features
|
Callable converting |
required | |
n
|
Number of training episodes. |
required | |
epsilon
|
Exploration rate or schedule for the epsilon-greedy policy. |
required | |
alpha
|
Step size or schedule. |
required | |
gamma
|
Discount factor in |
required | |
w
|
Initial action-value weight vector. If omitted, initialize it to zeros. |
None
|
|
max_episode_length
|
Maximum number of steps per episode. |
1000
|
|
verbose
|
Whether to print step-by-step diagnostics. |
False
|
|
history
|
Whether to return weights, returns, and episode lengths collected during training. |
False
|
|
rng
|
NumPy generator or integer seed for exploration and tie-breaking. |
None
|
Returns:
| Type | Description |
|---|---|
|
Learned weight vector. If |
Source code in gym_classics2/algorithms/linear_approximation.py
203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 | |
create_fourier_basis_coefs
Create Fourier basis coefficient vectors.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim
|
Number of state-feature dimensions. |
required | |
order
|
Maximum coefficient in each dimension. |
required |
Returns:
| Type | Description |
|---|---|
|
Array containing the Cartesian product of coefficients from zero through |
|
|
|
Source code in gym_classics2/algorithms/linear_approximation.py
transformation_fourier_basis
Create a Fourier basis feature transformation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
min
|
Minimum state value in each dimension. |
required | |
max
|
Maximum state value in each dimension. |
required | |
order
|
Maximum Fourier coefficient in each dimension. |
required |
Returns:
| Type | Description |
|---|---|
|
Callable that normalizes a state to the unit hypercube and returns its |
|
|
Fourier basis features. |
Example
Source code in gym_classics2/algorithms/linear_approximation.py
Eligibility traces
gym_classics2.algorithms.eligibility_traces
Semi-gradient Sarsa(lambda) with linear function approximation.
Callers provide a state_features(state, env) function that converts states
to feature vectors.
semi_gradient_Sarsa_lambda
semi_gradient_Sarsa_lambda(env, state_features, n, epsilon, alpha, gamma, lam, w=None, max_episode_length=1000, verbose=False, history=False, rng=None)
Run semi-gradient Sarsa(lambda) with linear function approximation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Episodic environment used to generate experience. |
required | |
state_features
|
Callable converting |
required | |
n
|
Number of episodes. |
required | |
epsilon
|
Exploration rate or schedule for the epsilon-greedy policy. |
required | |
alpha
|
Step size or schedule. |
required | |
gamma
|
Discount factor in |
required | |
lam
|
Trace-decay parameter lambda in |
required | |
w
|
Initial weight vector. If omitted, initialize it to zeros. |
None
|
|
max_episode_length
|
Maximum number of steps per episode. |
1000
|
|
verbose
|
Whether to print step-by-step diagnostics. |
False
|
|
history
|
Whether to return weights, returns, and episode lengths collected during training. |
False
|
|
rng
|
NumPy generator or integer seed for exploration and tie-breaking. |
None
|
Returns:
| Type | Description |
|---|---|
|
Learned weight vector. If |
Source code in gym_classics2/algorithms/eligibility_traces.py
18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 | |
Policy-gradient methods
gym_classics2.algorithms.policy_gradient_methods
This file implements policy gradient methods for learning parameterized policies. The main algorithm implemented is REINFORCE, which is a Monte Carlo policy gradient method that updates policy parameters based on the returns observed in sampled episodes. The policy is represented using a softmax function over linear state-action features, and the algorithm estimates the policy gradient using the log-likelihood of actions taken in the episodes. This implementation allows for learning stochastic policies that can handle exploration and exploitation in reinforcement learning tasks.
Callers provide a state_features(state, env) function that converts states
to feature vectors suitable for the environment being used.
h
Return the linear action preference for state s and action a.
pi
Return the softmax action-probability vector for a state.
Source code in gym_classics2/algorithms/policy_gradient_methods.py
sample_episode_approx_policy
Sample an episode from a parameterized policy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Episodic Gymnasium environment. |
required | |
pi
|
Callable returning action probabilities for |
required | |
theta
|
Policy parameter vector. |
required | |
max_episode_length
|
Maximum number of transitions to sample. |
1000
|
|
rng
|
NumPy generator or integer seed used to sample actions. |
None
|
Returns:
| Type | Description |
|---|---|
|
List of |
Source code in gym_classics2/algorithms/policy_gradient_methods.py
choose_action_w
Sample an action from a parameterized policy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Environment providing the discrete action space. |
required | |
pi
|
Callable returning action probabilities for |
required | |
theta
|
Policy parameter vector. |
required | |
state
|
Current state. |
required | |
rng
|
NumPy generator or integer seed used to sample the action. |
None
|
Returns:
| Type | Description |
|---|---|
|
Selected integer action ID. |
Source code in gym_classics2/algorithms/policy_gradient_methods.py
REINFORCE
REINFORCE(env, state_features, n, alpha, gamma, theta=None, max_episode_length=1000, verbose=False, history=False, rng=None)
Run REINFORCE with a linear softmax policy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Episodic environment used to generate experience. |
required | |
state_features
|
Callable converting |
required | |
n
|
Number of training episodes. |
required | |
alpha
|
Policy step size or schedule. |
required | |
gamma
|
Discount factor in |
required | |
theta
|
Initial policy parameter vector. If omitted, initialize it to zeros. |
None
|
|
max_episode_length
|
Maximum number of steps per episode. |
1000
|
|
verbose
|
Whether to print step-by-step diagnostics. |
False
|
|
history
|
Whether to return parameters, returns, and episode lengths collected during training. |
False
|
|
rng
|
NumPy generator or integer seed used to sample actions. |
None
|
Returns:
| Type | Description |
|---|---|
|
Learned policy parameters. If |
|
|
|
Source code in gym_classics2/algorithms/policy_gradient_methods.py
85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 | |
AC
AC(env, state_features, n, alpha_policy, alpha_value, gamma, max_episode_length=1000, verbose=False, history=False, rng=None)
Run one-step actor-critic with linear policy and value approximators.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
env
|
Episodic environment used to generate experience. |
required | |
state_features
|
Callable converting |
required | |
n
|
Number of training episodes. |
required | |
alpha_policy
|
Step size or schedule for policy updates. |
required | |
alpha_value
|
Step size or schedule for value-function updates. |
required | |
gamma
|
Discount factor in |
required | |
max_episode_length
|
Maximum number of steps per episode. |
1000
|
|
verbose
|
Whether to print step-by-step diagnostics. |
False
|
|
history
|
Whether to return policy parameters, returns, and episode lengths collected during training. |
False
|
|
rng
|
NumPy generator or integer seed used to sample actions. |
None
|
Returns:
| Type | Description |
|---|---|
|
Tuple |
|
|
If |
Source code in gym_classics2/algorithms/policy_gradient_methods.py
183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 | |
Schedules
gym_classics2.algorithms.schedules
Scalar schedules for reinforcement-learning hyperparameters.
A schedule is a callable that maps a nonnegative episode or step index t to
a floating-point value. The included algorithms use schedules to vary parameters
such as the step size (alpha) and exploration rate (epsilon) during training.
For example, a linear schedule can decrease epsilon from 1.0 to 0.1 over 10 steps:
schedule = LinearDecaySchedule(1.0, min_value=0.1, decay_steps=10)
for t in (0, 5, 10, 15):
print(t, schedule(t))
# 0 1.0
# 5 0.55
# 10 0.1
# 15 0.1
Schedule
Interface for a scalar value indexed by episode or time step.
Source code in gym_classics2/algorithms/schedules.py
__call__
Evaluate the schedule.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
t
|
Nonnegative episode or step index. |
required |
Returns:
| Type | Description |
|---|---|
|
Scheduled scalar value at |
ConstantSchedule
Bases: Schedule
Return the same value for every index.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Constant value returned by the schedule. |
required |
Source code in gym_classics2/algorithms/schedules.py
StepSchedule
Bases: Schedule
Switch once from a high value to a low value.
The schedule returns high_value while t < steps and low_value
from t == steps onward.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
high_value
|
Value before the switch. |
required | |
low_value
|
Value at and after the switch. |
required | |
steps
|
Index at which to switch to |
required |
Source code in gym_classics2/algorithms/schedules.py
LinearDecaySchedule
Bases: Schedule
Interpolate linearly from an initial value to a final value.
The schedule returns initial_value at t = 0, changes linearly through
decay_steps, and returns min_value thereafter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
initial_value
|
Value at index zero. |
required | |
min_value
|
Final value and lower endpoint of a decreasing schedule. |
required | |
decay_steps
|
Number of indices over which to interpolate. |
required |
Source code in gym_classics2/algorithms/schedules.py
ExponentialDecaySchedule
Bases: Schedule
Decay geometrically to a minimum value.
At index t, the value is max(min_value, initial_value * decay_rate**t).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
initial_value
|
Unclipped value at index zero. |
required | |
min_value
|
Lower bound for the returned value. |
required | |
decay_rate
|
Multiplicative factor applied at each successive index. |
required |
Source code in gym_classics2/algorithms/schedules.py
InverseDecaySchedule
Bases: Schedule
Decay in inverse proportion to the index.
The schedule returns initial_value at t = 0. For t > 0, it
returns max(min_value, initial_value / t).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
initial_value
|
Value at index zero and numerator of the inverse schedule. |
required | |
min_value
|
Lower bound for the returned value. |
0.0
|
Source code in gym_classics2/algorithms/schedules.py
plot_schedule
Plot scheduled values for indices 0 through steps - 1.
Example
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
schedule
|
Callable that accepts an integer index and returns a scalar. |
required | |
steps
|
Number of scheduled values to plot. |
1000
|