Skip to content
GenAI Learn/Reinforcement Learning for LLMs
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

Reinforcement Learning Foundations

8 min read

You'll learn to

  • -Explain the core RL loop: agent, environment, action, reward
  • -Compute a discounted return and explain why discounting matters
  • -Understand DQN's one-step TD target as a bootstrapped estimate

You met reinforcement learning briefly back in Phase 1, as one of four ways a system can learn. This module goes deep on it, for a specific reason: every reasoning-focused LLM released since 2024 - OpenAI's o1, DeepSeek-R1, and everything that has followed them - is trained with reinforcement learning on top of an already-trained base model. Understanding RL is no longer optional context for an AI engineer; it is directly how the most capable current models are made.

The Loop: Agent, Environment, Reward

An RL agent takes an action in some environment, the environment responds with a new state and a reward, and the agent uses that reward to get better at choosing actions next time. For an LLM being trained with RL, the "environment" is unusually abstract: a prompt is the state, generating a response is the action, and a reward model (or a rule-based checker, for something like verifiable math answers) scores how good that response was.

Discounted Return: What an Episode Actually Earned

An agent should not just care about its very next reward - it should care about its whole future, discounted so a reward far away matters less than one right now. This single number, the discounted return, is the quantity nearly every RL algorithm is ultimately trying to maximize.

Discounted Return
Gₜ = Σₖ γᵏ rₜ₊ₖ
Sum of every future reward, each one discounted by gamma raised to how many steps away it is.
Discounted return, from scratch

Bootstrapping: Learning Before an Episode Ends

Waiting for a full episode to finish before learning anything is slow and sometimes impossible (some tasks never naturally end). Deep Q-Networks (DQN) instead bootstrap from a one-step lookahead: this step's actual reward, plus a discounted estimate of the best the agent can do from here on, using the network's own current beliefs about the next state.

DQN's TD Target
target = r + γ · maxₐ Q(s', a)
This step's real reward, plus a discounted guess at the best achievable value from the next state - unless the episode just ended, in which case the target is just r.
The TD target Q-learning trains toward

AI Lab's Level 32 and 33 implement both of these directly: the discounted return formula, and the exact TD-target computation every DQN update trains toward.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement both from scratch: Level 55 (Discounted Returns) and Level 56 (Q-Values & the Bellman Update) in AI Lab.