TD(0)
TD(0) is a fundamental reinforcement learning algorithm that updates value estimates using immediate rewards and the estimated value of the next state.
106 plain-language definitions from the TiorAI glossary, filed under Reinforcement Learning. Every entry opens with a one-sentence definition, then explains where the term is used.
TD(0) is a fundamental reinforcement learning algorithm that updates value estimates using immediate rewards and the estimated value of the next state.
TD(lambda) is a reinforcement learning algorithm that combines temporal difference learning with eligibility traces to efficiently predict future rewards.
Temporal credit assignment is the process of determining which past actions or decisions led to current outcomes, especially when rewards or results are delayed over time.
Temporal Difference Learning is a reinforcement learning method where an agent learns to predict future rewards by updating estimates based on the difference between successive predictions over time.
Termination condition is a specific criterion that determines when a process or algorithm should stop executing.
Time Step is the discrete interval of time used in simulations and computational models to progress the system state incrementally.
Trajectory is the path that an object follows through space as a function of time.
Tree Policy is a set of guidelines or regulations that govern the planting, maintenance, and removal of trees in a specific area to ensure environmental sustainability and community safety.