Reinforcement Learning Foundations
You'll learn to
- -Explain the core RL loop: agent, environment, action, reward
- -Compute a discounted return and explain why discounting matters
- -Understand DQN's one-step TD target as a bootstrapped estimate
You met reinforcement learning briefly back in Phase 1, as one of four ways a system can learn. This module goes deep on it, for a specific reason: every reasoning-focused LLM released since 2024 - OpenAI's o1, DeepSeek-R1, and everything that has followed them - is trained with reinforcement learning on top of an already-trained base model. Understanding RL is no longer optional context for an AI engineer; it is directly how the most capable current models are made.
The Loop: Agent, Environment, Reward
An RL agent takes an action in some environment, the environment responds with a new state and a reward, and the agent uses that reward to get better at choosing actions next time. For an LLM being trained with RL, the "environment" is unusually abstract: a prompt is the state, generating a response is the action, and a reward model (or a rule-based checker, for something like verifiable math answers) scores how good that response was.
Discounted Return: What an Episode Actually Earned
An agent should not just care about its very next reward - it should care about its whole future, discounted so a reward far away matters less than one right now. This single number, the discounted return, is the quantity nearly every RL algorithm is ultimately trying to maximize.
Bootstrapping: Learning Before an Episode Ends
Waiting for a full episode to finish before learning anything is slow and sometimes impossible (some tasks never naturally end). Deep Q-Networks (DQN) instead bootstrap from a one-step lookahead: this step's actual reward, plus a discounted estimate of the best the agent can do from here on, using the network's own current beliefs about the next state.
AI Lab's Level 32 and 33 implement both of these directly: the discounted return formula, and the exact TD-target computation every DQN update trains toward.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement both from scratch: Level 55 (Discounted Returns) and Level 56 (Q-Values & the Bellman Update) in AI Lab.