td-lambda-interpolates-mc-and-td
IN premise — entries/2026/06/21/wiki-Reinforcement_learning-chunk-2.md
Created 2026-06-21T09:55:52+00:00
TD(λ) with λ=0 relies entirely on Bellman equations (pure TD), while λ=1 is equivalent to Monte Carlo with no Bellman reliance; λ provides continuous interpolation between the two