Environment API
Base environment
gym_classics2.envs.abstract.base_env.BaseEnv
Bases: Env
Base class for finite Gymnasium environments with explicit model access.
BaseEnv supplies seeded reset and step methods, a discrete action
space, optional integer encoding of states, and the complete transition model
returned by model. Subclasses define the environment dynamics through
_next_state, _reward, _done, and _generate_transitions. Each
generated transition has the form (next_state, reward, terminated,
probability).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
starts
|
Raw states from which an episode may start. |
required | |
action_labels
|
Labels whose positions define the integer action IDs. |
None
|
|
tabular
|
If true, observations are consecutive integer state IDs. If false, subclasses return raw states and must define a suitable observation space. |
True
|
|
reachable_states
|
All reachable raw states. If omitted, they are
discovered from |
None
|
Attributes:
| Name | Type | Description |
|---|---|---|
action_space |
Discrete Gymnasium action space. |
|
observation_space |
Discrete state space in tabular mode; otherwise defined by the subclass. |
|
state |
Current raw environment state, or |
|
tabular |
Whether observations use integer state IDs. |
Source code in gym_classics2/envs/abstract/base_env.py
9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 | |
start_states
property
Tuple of raw states from which an episode may start.
These are raw environment states even when :attr:tabular is true. Use
state2id to convert them to integer observations.
states
state2id
id2state
is_reachable
Returns True if the state can be reached from at least one start location, False otherwise.
actions
action2id
Converts a action label into a numeric action ID.
id2action
Converts a numeric action ID into a label. Choices for type are 'text' and 'arrow'.
Source code in gym_classics2/envs/abstract/base_env.py
model
Return the complete transition model for a state-action pair.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
Integer state ID when :attr: |
required | |
action
|
Integer action ID. |
required |
Returns:
| Type | Description |
|---|---|
|
A four-item list |
|
|
|
|
|
otherwise. The remaining items are one-dimensional NumPy arrays in the |
|
|
same order. |
Source code in gym_classics2/envs/abstract/base_env.py
Linear walk
gym_classics2.envs.abstract.linear_walk.LinearWalk
Bases: BaseEnv
Finite one-dimensional walk with terminal outcomes at both ends.
Raw states are integer positions from 0 through length - 1. Every
episode starts at the center position. Action 0 moves left and action 1
moves right; movement is deterministic. Attempting to move left from position
0 or right from position length - 1 terminates the episode and yields
the corresponding boundary reward. All other transitions have reward zero.
Observations are consecutive integer state IDs. For this environment, each ID is equal to its raw position.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
length
|
Odd number of nonterminal positions in the walk. |
required | |
left_reward
|
Reward for terminating beyond the left boundary. |
required | |
right_reward
|
Reward for terminating beyond the right boundary. |
required |
Source code in gym_classics2/envs/abstract/linear_walk.py
Gridworld
gym_classics2.envs.abstract.gridworld.Gridworld
Bases: BaseEnv
Finite rectangular gridworld constructed from an ASCII layout.
Layout rows are written from top to bottom; raw states are (x, y)
coordinates with (0, 0) in the lower-left corner. S marks a start,
G a terminal goal, X a blocked cell, and a space a traversable cell.
Other characters are traversable labels retained for plotting. Optional |
characters are ignored when the layout is parsed.
Actions and transitions:
The default actions move up, right, down, and left and are represented in that
order by the integers 0 to 3. The default transition model is deterministic.
An action that would leave the gridworld or enter
a blocked cell leaves the agent in place. Additional actions can be added in the constructor and will use
integer IDs starting at 4. The transition model can be changed
by subclassing and overwriting _next_state.
For an example of a stochastic transition model, see ClassicGridworld.
Reward model:
Entering a goal yields goal_reward and terminates the episode; other transitions yield
step_reward. Since goal states are terminal (absorbing), the reward for a transition from
a goal state is always zero.
The default setting is for a world without step cost and reaching the goal is rewarded with 1.0. If you
want to use a step cost, set step_reward to a negative value. Note that the goal_reward
is for the transition to the goal state and thus needs to be
adjusted to reflect the positive reward for reaching the goal minus the cost of getting there.
You can also overwrite the _reward method to implement a custom reward model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
layout_string
|
Rectangular ASCII representation of the grid. |
required | |
action_labels
|
Labels for the four actions, ordered as up, right, down, and left. |
['up', 'right', 'down', 'left']
|
|
goal_reward
|
Reward for a transition into a goal cell. |
1.0
|
|
step_reward
|
Reward for any other transition. |
0.0
|
|
tabular
|
If true, observations are integer state IDs. If false, they are
raw |
True
|
|
render_mode
|
|
None
|
Note
Rendering requires the optional render dependency extra.
Source code in gym_classics2/envs/abstract/gridworld.py
28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 | |
reset
Reset to a start cell and render a frame in human mode.
Source code in gym_classics2/envs/abstract/gridworld.py
step
Advance the environment and render a frame in human mode.
Source code in gym_classics2/envs/abstract/gridworld.py
render
Prints a gridworld array in a human-readable format. The array should be a vector with values for states in the gridworld, such as a value function or policy.
Source code in gym_classics2/envs/abstract/gridworld.py
image
image(V=None, policy=None, episode=None, labels=None, title=None, cmap='auto', origin='lower', clim=None)
Display the gridworld, values, and optional actions as an image.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
V
|
One value per state, such as a value function. If omitted, cells display their state IDs and the color bar is hidden. |
None
|
|
policy
|
One action ID per state. Actions are drawn as arrows and
replace |
None
|
|
episode
|
Sequence of transitions whose first two entries are the state
and action IDs. Actions taken in visited states are drawn as arrows
and replace |
None
|
|
labels
|
One label per state. If |
None
|
|
title
|
Plot title. |
None
|
|
cmap
|
Matplotlib colormap or colormap name. |
'auto'
|
|
origin
|
|
'lower'
|
|
clim
|
Optional |
None
|
Source code in gym_classics2/envs/abstract/gridworld.py
image_list
Creates a sequence of images, one for each episode.
Source code in gym_classics2/envs/abstract/gridworld.py
Noisy gridworld
gym_classics2.envs.abstract.noisy_gridworld.NoisyGridworld
Bases: Gridworld
Gridworld with classic 80-10-10 stochastic action outcomes.
The requested action is executed with probability 0.8. With probability 0.1
each, it is instead rotated 90 degrees clockwise or counterclockwise. The
resulting move follows the boundary, blocking, reward, and termination rules
defined by Gridworld. The model method enumerates all three
possible outcomes, including duplicate next states when movement is blocked.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
layout_string
|
Rectangular ASCII representation of the grid. |
required | |
action_labels
|
Labels for the four actions, ordered as up, right, down, and left. |
['up', 'right', 'down', 'left']
|
|
goal_reward
|
Reward for a transition into a goal cell. |
1.0
|
|
step_reward
|
Reward for any other transition. |
0.0
|
|
tabular
|
If true, observations are integer state IDs. If false, they are
raw |
True
|
|
render_mode
|
|
None
|
Note
Rendering requires the optional render dependency extra.
Source code in gym_classics2/envs/abstract/noisy_gridworld.py
Concrete environments
gym_classics2.envs.gym_classics2.classic_gridworld_v1.ClassicGridworld
Bases: Gridworld
A 4x3 pedagogical gridworld. The agent starts in the bottom-left cell. Actions are noisy; with a 10% chance each, a move action may be rotated by 90 degrees clockwise or counter-clockwise (the "80-10-10 rule"). Cell (1, 1) is blocked and cannot be occupied by the agent.
Reference: Russell and Norvig, Artificial Intelligence: A Modern Approach (3rd ed., 2010), p. 646.
state: Grid location.
actions: Move up/right/down/left.
rewards: +1 for taking any action in cell (3, 2). -1 for taking any action in cell (3, 1). NOTE: v1 uses the original -0.04 penalty for each state.
termination: Earning a nonzero reward.
Source code in gym_classics2/envs/gym_classics2/classic_gridworld_v1.py
gym_classics2.envs.gym_classics2.cliff_walk_v1.CliffWalk
Bases: Gridworld
The Cliff Walking task, a 12x4 gridworld often used to contrast Sarsa with Q-Learning. The agent begins in the bottom-left cell and must navigate to the goal (bottom-right cell) without entering the region along the bottom ("The Cliff").
v1 follows the textbook and does not end episodes when the cliff is reached. Also, the goal is a real state.
Reference: Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., 2018), p. 132, Example 6.6.
state: Grid location.
actions: Move up/right/down/left.
rewards: -100 for entering The Cliff. -1 for all other transitions.
termination: reaching the goal.
Source code in gym_classics2/envs/gym_classics2/cliff_walk_v1.py
gym_classics2.envs.gym_classics2.dyna_maze.DynaMaze
Bases: Gridworld
A 9x6 deterministic gridworld with barriers to make navigation more challenging. The agent starts in cell (0, 3); the goal is the top-right cell.
Reference: Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., 2018), p. 164, Example 8.1.
state: Grid location.
actions: Move up/right/down/left.
rewards: +1 for episode termination.
termination: Reaching the goal.
Source code in gym_classics2/envs/gym_classics2/dyna_maze.py
gym_classics2.envs.gym_classics2.four_rooms.FourRooms
Bases: Gridworld
An 11x11 gridworld segmented into four rooms. The agent begins in the bottom-left cell; the goal is in the top-right cell.
Reference: Sutton, Precup, and Singh, Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning (1999), p. 192.
state: Grid location.
actions: Move up/right/down/left.
rewards: +1 for episode termination.
termination: Taking any action in the goal.
Source code in gym_classics2/envs/gym_classics2/four_rooms.py
gym_classics2.envs.gym_classics2.L_maze.LMazeGridworld
Bases: Gridworld
A deterministic 10x10 maze separated by an L-shaped barrier.
The agent begins below the horizontal barrier and must travel around it to reach the goal near the upper-right corner.
Reference: Dijkstra's algorithm, Wikipedia.
Source code in gym_classics2/envs/gym_classics2/L_maze.py
gym_classics2.envs.gym_classics2.sparse_gridworld.SparseGridworld
Bases: NoisyGridworld
A 10x8 featureless gridworld. The agent starts in cell (1, 3) and the goal is at
cell (6, 3). To make it more challenging, the same 80-10-10 transition probabilities
from ClassicGridworld are used. Great for testing various forms of credit
assignment in the presence of noise.
Reference: Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., 2018), p. 147, Figure 7.4.
states: Grid location.
actions: Move up/right/down/left.
rewards: +1 for episode termination.
termination: Reaching the goal.
Source code in gym_classics2/envs/gym_classics2/sparse_gridworld.py
gym_classics2.envs.gym_classics2.windy_gridworld.WindyGridworld
Bases: Gridworld
A 10x7 deterministic gridworld where some columns are affected by an upward wind. The agent starts in cell (0, 3) and the goal is at cell (7, 3). If an agent executes an action from a cell with wind, the resulting position is given by the vector sum of the action's effect and the wind.
Reference: Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., 2018), p. 130, Example 6.5.
state: Grid location.
actions: Move up/right/down/left.
rewards: -1 for all transitions unless the episode terminates.
termination: Reaching the goal.
Source code in gym_classics2/envs/gym_classics2/windy_gridworld.py
gym_classics2.envs.gym_classics2.linear_walks.Walk5
Bases: LinearWalk
A 5-state deterministic linear walk. Ideal for implementing random walk experiments.
Reference: Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., 2018), p. 125, Example 6.2.
state: Discrete position {0, ..., 4} on the number line.
actions: Move left/right.
rewards: +1 for moving right in the extreme right state.
termination: Moving right in the extreme right state or moving left in the extreme left state.
Source code in gym_classics2/envs/gym_classics2/linear_walks.py
gym_classics2.envs.gym_classics2.linear_walks.Walk19
Bases: LinearWalk
Same as 5Walk but with 19 states and an additional -1 reward for moving left
in the extreme left state.
Reference: Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., 2018), p. 145, Example 7.1.