Glossary
Post-training
Post-training is everything done to shape a model after its main training run, and it is where almost all of the behaviour a user notices actually comes from.
In plain terms
The shaping that happens after the expensive learning is finished. The first stage decides what it knows; this decides how it acts, what it declines, how it lays an answer out and what it sounds like. Almost everything anybody notices when using one of these things was decided here rather than earlier.
Why it matters
Because it explains why two products built on similar foundations feel nothing alike. Tone, willingness, formatting habits and where the refusals sit are all decided at this stage, so a comparison between assistants is mostly a comparison of shaping choices rather than of underlying knowledge, and buyers who assume otherwise pick on the wrong axis.
How it works
The division of labour is clean and worth holding onto. The main run decides what the model knows and how well it can use language, while this stage decides how it behaves with that capability, which is why shaping can make something markedly more useful without teaching it a single new fact.
Human preference is the usual instrument. People compare candidate answers, those comparisons train a stand-in that can score at a scale no group could reach, and the model is then shaped towards what that stand-in rewards, which is how a general capability becomes a particular manner. The choices embedded there are editorial rather than technical.
Refusal behaviour is manufactured here rather than being an inherent property. What a system declines, how firmly and with what explanation are all shaping decisions, which is why the same underlying capability can be permissive in one product and cautious in another and why those positions move between versions.
It is where most of the visible improvement between releases comes from. Successive versions frequently feel substantially better at following instructions, laying out answers and handling awkward requests while knowing much the same things, because that is the part being worked on continuously and the expensive part is not repeated often.
Which stage decided what you notice
Seen in the wild
An open-weight line published for self-hosting, where the shaping decisions arrive with the weights and can be examined rather than inferred.
QwenA price disruptor whose successive generations stayed competitive on reasoning and coding rather than only on cost, which is shaping work rather than scale.
DeepSeekAn assistant whose distinguishing strengths are described as understanding detailed instructions and holding context, both of which are manners rather than knowledge.
Claude
Common misconceptions
People assume
It teaches the model new knowledge.
In fact
It teaches behaviour. What the model knows is largely settled by the earlier expensive run, and this stage decides what it does with that, which is why shaping can transform how useful something feels without changing a single fact it holds.
People assume
Refusals are a property of the model.
In fact
They are a manufactured position, chosen at this stage and adjustable. That is why the same underlying capability is permissive in one product and cautious in another, and why a system's willingness can change noticeably between versions without anything about its knowledge changing.
Telling them apart
Post-training vs Pre-training
Post-training
How it behaves: manners, format, refusals.
What it knows, and how well it handles language.
The second is enormously expensive and done rarely; the first is comparatively cheap and done continuously, which is why releases improve manners faster than knowledge.
Questions
- Why do two assistants on similar foundations feel so different?
- Because almost everything you notice was decided at this stage. Tone, how willingly it attempts something, how it lays out an answer and where it draws its lines are all shaping choices, so comparing assistants is largely comparing editorial decisions rather than underlying knowledge. That is a fairer description of what you are choosing between.
- Is this the same as adapting a model to our own material?
- It is the same family of technique applied by the maker at a different scale and for a different purpose. They are shaping general behaviour for everybody; an organisation adapting a model afterwards is shaping it towards one house style. The mechanisms overlap and the intent and the scale do not.
- Does it explain why a new version feels better?
- Usually yes, and more than any change in knowledge does. The expensive run happens rarely while shaping continues, so successive releases commonly follow instructions better, format more helpfully and handle awkward requests more gracefully while knowing much the same things. Most of what a user calls improvement lives here.
Key takeaways
- The earlier run decides knowledge; this decides behaviour.
- Human preference is the instrument, which makes the choices editorial.
- Refusals are manufactured here, not inherent, and they move between versions.
- Most visible improvement between releases comes from this stage.
Last checked July 2026