ppo-trpo-a3c-on-policy-advantage

IN premiseentries/2026/06/21/wiki-Reinforcement_learning-chunk-3.md

Created 2026-06-21T09:55:52+00:00

PPO, TRPO, and A3C are on-policy algorithms that use the advantage function, while SAC, TD3, and DDPG are off-policy algorithms