AI Fundamentals

Reasoning and Reinforcement Learning: Beyond Pretraining

AI Foundations #21 explains reasoning tokens, inference-time compute, reward signals, RLHF, PPO, and how reinforcement-learning post-training changes model behavior.

Approximately 14 min read · AI Foundations / Lesson 21

Yesterday’s lesson on multimodal AI changed the kind of information a model can receive.

Text can enter through tokens.

Images can enter through visual features.

Audio can enter through an audio encoder.

After all of that information reaches the model, however, a familiar question remains:

How does the model decide what to do with it, and why do some modern models spend many tokens working through a problem before giving the final answer?

That leads to two ideas that are often discussed together but should not be confused:

reasoning at inference time and reinforcement learning during training.

When I first started reading about reasoning models, I imagined that somebody had added a new “reasoning engine” beside the transformer.

That is not the mental model I use now.

A better starting point is:

pretrained transformer
+ post-training
+ prompts and tools
+ more inference-time tokens
= behavior that can look much more deliberate

The same basic transformer machinery is still there.

The weights are still numbers.

The forward pass still produces logits.

Sampling still chooses tokens.

The KV cache still stores earlier attention state.

What changes is how the model has been trained to use those mechanisms and how much computation it is allowed to spend before answering.

Reasoning is not a separate mathematical organ

Suppose I ask a model:

A store discounts a $120 item by 25%, then charges 13% tax.
What is the final price?

A model could jump directly to an answer.

Or it could generate intermediate steps:

25% of 120 = 30
discounted price = 90
13% tax on 90 = 11.70
final price = 101.70

From the outside, the second response looks more like reasoning.

Inside the autoregressive model, something very ordinary is happening:

prompt
-> predict next token
-> append token
-> predict next token
-> append token
-> repeat

The intermediate tokens become part of the context for later tokens.

That means each extra step gives the model another opportunity to transform the problem into a form that is easier to continue from.

This was my first useful mental model:

A reasoning trace can act like a scratchpad made out of tokens.

The model is not pausing the transformer and switching to a completely different algorithm.

It is repeatedly running the same kind of next-token computation over a growing context.

More reasoning tokens mean more inference-time compute

This connects directly to our lessons on inference, sampling, and KV cache.

Imagine two responses to the same question.

Response A:

42

Response B:

First I identify the quantities...
Then I compute...
Now I check...
Therefore the answer is 42.

Even if both use the same model weights, Response B requires more decode steps.

Roughly:

more generated tokens
-> more forward passes
-> more KV-cache growth
-> more latency
-> more compute

So “reasoning harder” can partly mean spending a larger inference-time token budget.

That is one reason reasoning models can be slower even when the underlying parameter count is unchanged.

The tradeoff is not mysterious:

less test-time compute -> faster, cheaper answer
more test-time compute -> more opportunities to decompose and check

More tokens do not guarantee a better answer.

They simply give the model more computational steps in which a better answer may emerge.

A written explanation is not automatically a faithful window into the model

This distinction matters.

If a model prints:

I chose answer B because...

we should not automatically assume that every sentence is a perfect record of the internal causal process that produced B.

The output is generated text.

It can be useful.

It can expose arithmetic mistakes.

It can give us a structure to verify.

But it can also contain a plausible-sounding explanation that is incomplete or wrong.

So I separate two claims:

claim 1:
intermediate generated tokens can help solve some tasks

claim 2:
the generated explanation perfectly reveals the model's internal reasoning

The first can be useful without the second being true.

That keeps “reasoning” from turning into a magical word.

Where does reinforcement learning enter?

Pretraining taught us the basic objective:

given previous tokens
-> predict the next token

The model sees an enormous amount of data and adjusts its weights to reduce prediction loss.

After pretraining, we may want something more specific than “continue text plausibly.”

We may want behavior such as:

follow the user's instruction
prefer a helpful answer over an unhelpful one
solve a math problem correctly
produce code that passes tests
avoid certain failure modes
use a desired response format

This is where post-training becomes important.

One family of post-training methods uses reinforcement learning, or RL.

Instead of learning only from a target next token, the system can learn from a reward signal that scores an action or completed response.

My simplest RL mental model

I started with four pieces:

state
action
reward
policy

A classic example might be a game.

The state is the current board.

The action is a move.

The reward tells us whether that move eventually helped win.

The policy is the strategy that maps states to action probabilities.

For a language model, the mapping is less visually obvious, but the idea carries over.

We can think conceptually about:

state:
prompt + generated tokens so far

action:
choose the next token, or more practically a sequence of tokens

policy:
the language model's probability distribution over possible outputs

reward:
a score for the resulting behavior

The model is already a probabilistic policy because it produces probabilities over tokens.

RL gives us a way to update that policy using rewards.

A reward is not the same thing as the training loss

This confused me at first because both eventually influence parameter updates.

A reward is usually easier to think about as the outcome we want:

high reward -> desirable behavior
low reward -> undesirable behavior

The training algorithm then turns that signal into an objective that can be optimized.

At the bottom of the stack, our earlier lessons return:

objective
-> gradients
-> backpropagation
-> weight updates

So RL does not throw away gradient descent.

It changes where the learning signal comes from and how the objective is constructed.

The parameters are still updated through math we have already studied.

RLHF: reinforcement learning from human feedback

One famous post-training pipeline is RLHF.

A simplified version looks like this:

1. start with a pretrained model
2. supervised fine-tune it on demonstrations
3. collect human preferences between candidate answers
4. train a reward model from those preferences
5. optimize the language-model policy against the reward model

The InstructGPT paper is a well-known example of this pattern.

Suppose a person sees two answers:

Answer A
Answer B

and prefers A.

That does not directly tell the model the exact correct token for every position.

Instead, many preference comparisons can train a model that estimates which responses people tend to prefer.

That reward model can then score new outputs.

The language-model policy is updated to increase expected reward while usually being constrained from moving too far from a reference policy.

I think of this as changing the question from:

"What token usually comes next in this dataset?"

toward:

"What kind of response earns a better score under the post-training objective?"

It is still language modeling, but the optimization target has changed.

Why PPO appears so often

One algorithm historically associated with RLHF is Proximal Policy Optimization, or PPO.

The full derivation is beyond what I needed for my first mental model.

The important idea is control.

Suppose the old policy assigned some probability to an action and the updated policy wants to assign a very different probability.

A rough ratio is:

r = probability under new policy
    ----------------------------
    probability under old policy

If training blindly rewards every large change, one update can move the policy too aggressively.

PPO uses a clipped objective designed to discourage updates from moving too far in one step.

My non-CS interpretation is similar to changing a financial forecast model.

If one quarter’s new data says a parameter should move, I do not necessarily want a single observation to multiply that parameter by ten overnight.

I want the update to respond to evidence without making the whole system unstable.

PPO’s clipping idea is one mechanism for keeping policy updates in a more controlled range.

The reference model is another kind of guardrail

Many RLHF-style systems also use a penalty related to how far the new policy moves from a reference model.

One common tool is KL divergence.

Earlier in this course we did not need the full formula, so I still do not need it here.

The intuition is enough:

new policy very similar to reference -> small divergence
new policy very different -> larger divergence

If reward optimization is the accelerator, a KL-style penalty can act like resistance against changing the model too far, too quickly.

Why might that matter?

Because a reward model is not perfect.

If the policy discovers a strange way to exploit the scorer, maximizing reward alone may produce behavior that looks nothing like what we intended.

This problem is often described as reward hacking.

The reward is a training signal, not a perfect definition of intelligence or truth.

Verifiable rewards change the game

Human preference is expensive and subjective.

Some tasks have something much cleaner.

For math, the final answer may be automatically checked.

For code, unit tests can run.

For formal problems, a verifier may decide whether the result satisfies the rules.

That gives a reward source such as:

correct final answer -> positive reward
wrong final answer   -> low or zero reward

This is attractive because the feedback can be generated at much larger scale without asking a human to compare every pair of outputs.

DeepSeek-R1 is an important recent example in the reasoning discussion. Its paper reports that reinforcement learning with verifiable signals can elicit behaviors such as longer problem solving, self-checking, and strategy changes.

I read that carefully.

It does not mean:

RL automatically creates truth

It means that if the training environment can reward successful problem solving, the policy can learn output strategies that are more likely to receive that reward.

The reward design still matters.

The task distribution still matters.

The verifier can still have blind spots.

Why longer reasoning can emerge from RL

Imagine two policies attempting a hard math problem.

Policy A often produces:

quick guess -> answer

Policy B often produces:

decompose
-> calculate
-> notice inconsistency
-> revise
-> answer

If the second pattern succeeds more often on tasks with reliable rewards, training can increase the probability of behavior resembling Policy B.

Nothing requires us to create a special “self-correction neuron.”

The optimization can shift probabilities across many tokens and many internal states so that longer solution trajectories become more likely.

This connects RL directly back to the first lesson in the whole series:

behavior changes
because weights change

The architecture may remain almost identical while post-training changes which trajectories through token space are likely.

Inference-time reasoning and RL are different operations

This distinction is worth making explicit.

Reinforcement-learning post-training

happens during training
computes an optimization objective
uses gradients / backpropagation
changes model weights

Reasoning at inference time

happens after training
runs forward passes
generates more tokens
does not normally backpropagate
does not normally update weights

A model can therefore have been trained with RL and then spend a variable amount of time reasoning at inference.

Those are two different axes:

how was the model trained?
how much compute does it spend answering this prompt?

I used to merge them into one idea because both appear in the phrase “reasoning model.”

Keeping them separate makes the system much easier to understand.

Sampling still matters

A reasoning-tuned model still eventually faces a probability distribution over next tokens.

That means our Sampling lesson did not become obsolete.

Temperature, top-p, top-k, and deterministic decoding can still alter the trajectory.

For a long solution, an early token choice can change everything that follows because that token becomes part of the next context.

Conceptually:

step 1 chooses token A instead of B
-> context changes
-> step 2 logits change
-> later reasoning path changes
-> final answer can change

This is one reason evaluating reasoning models requires more than looking at one pretty example.

The generation policy is stochastic unless decoding is made deterministic.

Context and KV cache matter even more

Long reasoning chains also reconnect to AI Foundations #17.

Every generated token extends the sequence.

The runtime may need to preserve KV state for those tokens.

So a model that produces 8,000 reasoning tokens before a 200-token answer has very different serving economics from a model that produces a 200-token answer immediately.

The cost shows up as:

decode time
KV-cache memory
context consumption
server occupancy
user-visible latency

This is why “the model supports a huge context” and “the model can reason for a long time cheaply” are not the same claim.

Context capacity is a limit.

Reasoning length is a workload.

Quantization and MoE are still underneath

Nothing about reasoning removes the systems topics we learned later.

A reasoning model can still be quantized.

It can still use Mixture of Experts.

It can still be multimodal.

A real stack might look like:

multimodal input
-> embeddings / modality encoders
-> MoE transformer
-> quantized weights at inference
-> KV cache
-> long reasoning generation
-> sampled answer

That is why the Foundations sequence was cumulative.

Each new topic adds another layer without erasing the previous ones.

A practical example: math reward versus preference reward

Suppose a model answers:

What is 37 x 24?

There are different ways to train behavior around this task.

With supervised fine-tuning, the dataset might provide a demonstration containing a worked solution.

With human preference training, people might compare two candidate explanations and choose the clearer or more correct one.

With a verifiable reward, a system can parse the final number and check whether it equals:

888

Each signal teaches something slightly different.

A preference score can include style, helpfulness, and presentation.

A verifier can provide a very crisp correctness signal.

Neither one alone necessarily captures every property we care about.

That is why modern post-training systems can combine multiple stages and multiple objectives rather than relying on a single magic reward.

What RL cannot guarantee

It is tempting to summarize the whole subject as:

pretraining creates knowledge
RL creates reasoning

That is too simple.

Pretrained models can already show reasoning-like behavior.

Fine-tuning can improve task performance without RL.

RL can improve some behaviors while harming others if the objective is poor.

A model can generate a long chain and still reach the wrong answer.

A reward can be exploited.

A verifier can be incomplete.

A preference model can encode the preferences and biases of its data.

So the careful statement is:

Reinforcement learning is one set of methods for changing a model’s behavior using reward signals, and it can be especially useful when success can be evaluated more directly than by next-token imitation alone.

That is powerful enough without turning it into mythology.

The entire Foundations path in one picture

We started with the smallest pieces:

parameters
weights
biases

Then we learned how they are organized and used:

tensors
-> forward pass
-> training
-> loss
-> gradient descent
-> backpropagation

Then we moved into language:

embeddings
-> tokenization
-> attention
-> transformer

Then model life cycle:

pretraining
-> fine-tuning
-> inference
-> sampling

Then runtime systems:

context / KV cache
-> quantization
-> Mixture of Experts

Then richer inputs:

multimodal AI

And now the final layer in the core sequence:

reasoning behavior
+ reinforcement-learning post-training

When I put it all together, a modern reasoning model no longer looks like one mysterious object.

It looks like a stack of understandable mechanisms.

My final mental model

Here is the picture I would have wanted when this series began:

data
-> tokens or multimodal representations
-> embeddings
-> transformer forward passes
-> logits
-> token choices
-> generated sequence

Before deployment, training can change the weights through:

prediction losses
supervised examples
preference signals
verifiable rewards
-> objective
-> gradients
-> backpropagation
-> parameter updates

At inference time, the trained policy can spend more or fewer generated tokens working through a problem.

Those tokens become context.

That context affects later logits.

The model can check, revise, or continue because each new token creates another computational step.

There is no need to insert a mysterious box labeled “intelligence.”

The mechanisms we learned across the previous twenty lessons are enough to build a useful first-principles picture.

Where to go next

This lesson closes the core AI Foundations sequence.

That does not mean the subject is finished. It means the foundation is finally strong enough to branch into deeper tracks without treating every new acronym as a separate mystery.

From here, the natural directions are:

But the map underneath them stays the same:

weights store what training changed; forward passes use those weights; token generation creates the trajectory; and training objectives decide which trajectories become more likely.

Sources and further reading

Continue reading