rl-environment-modeled-as-mdp
IN premise — entries/2026/06/21/wiki-Reinforcement_learning-chunk-1.md
Created 2026-06-21T09:55:52+00:00
The RL environment is typically modeled as a Markov decision process (MDP) with four components: state space S, action space A, transition probability P, and reward function R