Skip to content

Glossary

Post-training

Post-training is everything done to shape a model after its main training run, and it is where almost all of the behaviour a user notices actually comes from.

In plain terms

The shaping that happens after the expensive learning is finished. The first stage decides what it knows; this decides how it acts, what it declines, how it lays an answer out and what it sounds like. Almost everything anybody notices when using one of these things was decided here rather than earlier.

01

Why it matters

Because it explains why two products built on similar foundations feel nothing alike. Tone, willingness, formatting habits and where the refusals sit are all decided at this stage, so a comparison between assistants is mostly a comparison of shaping choices rather than of underlying knowledge, and buyers who assume otherwise pick on the wrong axis.

02

How it works

The division of labour is clean and worth holding onto. The main run decides what the model knows and how well it can use language, while this stage decides how it behaves with that capability, which is why shaping can make something markedly more useful without teaching it a single new fact.

Human preference is the usual instrument. People compare candidate answers, those comparisons train a stand-in that can score at a scale no group could reach, and the model is then shaped towards what that stand-in rewards, which is how a general capability becomes a particular manner. The choices embedded there are editorial rather than technical.

Refusal behaviour is manufactured here rather than being an inherent property. What a system declines, how firmly and with what explanation are all shaping decisions, which is why the same underlying capability can be permissive in one product and cautious in another and why those positions move between versions.

It is where most of the visible improvement between releases comes from. Successive versions frequently feel substantially better at following instructions, laying out answers and handling awkward requests while knowing much the same things, because that is the part being worked on continuously and the expensive part is not repeated often.

Which stage decided what you notice

Which stage decided what you noticeThe practical consequence of this split is that a great deal of assistant comparison is measuring the wrong thing sincerely. Two products can sit on foundations of similar capability and produce experiences a user would describe as completely different: one that attempts almost anything and formats loosely, one that is careful and structured, one that refuses a category of request the other handles without comment. None of those differences is about knowledge and all of them are about editorial choices made after the expensive part finished. That is why a team's preference between assistants is so often stable and hard to justify on benchmarks, and why it is a legitimate basis for choosing rather than a failure of rigour. What is being preferred is a set of manners, chosen deliberately by people, and manners are a reasonable thing to have opinions about.The expensive runWhat it knows.How well it handles language.Done rarely, at great cost.Nearly invisible to a user.The shapingTone and willingness.How an answer is laid out.What it declines, and how.Almost everything you notice.Buyers compare products and arealmost entirely comparing theright column, while thediscourse is almost entirelyabout the left one.
The practical consequence of this split is that a great deal of assistant comparison is measuring the wrong thing sincerely. Two products can sit on foundations of similar capability and produce experiences a user would describe as completely different: one that attempts almost anything and formats loosely, one that is careful and structured, one that refuses a category of request the other handles without comment. None of those differences is about knowledge and all of them are about editorial choices made after the expensive part finished. That is why a team's preference between assistants is so often stable and hard to justify on benchmarks, and why it is a legitimate basis for choosing rather than a failure of rigour. What is being preferred is a set of manners, chosen deliberately by people, and manners are a reasonable thing to have opinions about.
03

Seen in the wild

  • An open-weight line published for self-hosting, where the shaping decisions arrive with the weights and can be examined rather than inferred.

    Qwen
  • A price disruptor whose successive generations stayed competitive on reasoning and coding rather than only on cost, which is shaping work rather than scale.

    DeepSeek
  • An assistant whose distinguishing strengths are described as understanding detailed instructions and holding context, both of which are manners rather than knowledge.

    Claude
04

Common misconceptions

People assume

It teaches the model new knowledge.

In fact

It teaches behaviour. What the model knows is largely settled by the earlier expensive run, and this stage decides what it does with that, which is why shaping can transform how useful something feels without changing a single fact it holds.

People assume

Refusals are a property of the model.

In fact

They are a manufactured position, chosen at this stage and adjustable. That is why the same underlying capability is permissive in one product and cautious in another, and why a system's willingness can change noticeably between versions without anything about its knowledge changing.

05

Telling them apart

Post-training vs Pre-training

Post-training

How it behaves: manners, format, refusals.

Pre-training

What it knows, and how well it handles language.

The second is enormously expensive and done rarely; the first is comparatively cheap and done continuously, which is why releases improve manners faster than knowledge.

06

Questions

Why do two assistants on similar foundations feel so different?
Because almost everything you notice was decided at this stage. Tone, how willingly it attempts something, how it lays out an answer and where it draws its lines are all shaping choices, so comparing assistants is largely comparing editorial decisions rather than underlying knowledge. That is a fairer description of what you are choosing between.
Is this the same as adapting a model to our own material?
It is the same family of technique applied by the maker at a different scale and for a different purpose. They are shaping general behaviour for everybody; an organisation adapting a model afterwards is shaping it towards one house style. The mechanisms overlap and the intent and the scale do not.
Does it explain why a new version feels better?
Usually yes, and more than any change in knowledge does. The expensive run happens rarely while shaping continues, so successive releases commonly follow instructions better, format more helpfully and handle awkward requests more gracefully while knowing much the same things. Most of what a user calls improvement lives here.
07

Key takeaways

  • The earlier run decides knowledge; this decides behaviour.
  • Human preference is the instrument, which makes the choices editorial.
  • Refusals are manufactured here, not inherent, and they move between versions.
  • Most visible improvement between releases comes from this stage.
09

Tools that use this

  • Qwen

    Open weights, so shaping arrives examinable rather than inferred.

  • DeepSeek

    Generations competitive on reasoning, which is shaping not scale.

  • Claude

    Instruction-following and context named as distinguishing strengths.

Last checked July 2026

All glossary terms