Skip to content

Glossary

Context degradation

The tendency for material in the middle of a very long input to be used less reliably than material at either end of it.

In plain terms

Give a system a very long document and it uses the beginning and the end more dependably than the middle. Room to put something is not the same as attention paid to it, and the two get sold as one number.

01

Why it matters

Because the headline capacity is what gets compared between tools and the reliable capacity is what your work depends on. A team that fills the available room because it is available has quietly moved the material it cares about into the least dependable position.

02

How it works

Capacity and reliability are different properties. The published figure describes how much can be sent, and nothing about it promises that everything sent will be used with equal care, which is the assumption most people make from it.

The middle is where the effect concentrates. Material at the start and end tends to be handled more dependably, so the position of a fact within a long input changes how likely it is to be used, independent of how relevant it is.

The failure is silent and looks like ignorance. Nothing reports that something in the middle was passed over, so the output simply omits it, and the natural conclusion is that the model did not know rather than that it did not reach.

It gets worse as the input grows, which is the opposite of what a larger window implies. Filling more of the available room puts more material into the region where reliability is lowest, so using the whole capacity is a decision with a cost.

Sending less is the fix that works. Retrieving the few passages that matter and sending those beats sending everything and relying on the system to find them, and it is cheaper as well, which makes it an unusually easy trade.

Where everything genuinely must go in, position is a lever you control. Putting what matters most at the beginning or the end, and restating a critical instruction at both, costs nothing and works with the effect rather than against it.

Where a fact sits in a long input

Where a fact sits in a long inputThe shape is what makes this counterintuitive, because most degradation people meet is monotonic: things get worse the further you go, and you compensate by putting the important part first. Here the far end recovers, which means the intuition of front-loading everything is only half right and the material at genuine risk is whatever sits in the region nobody thinks about. It also explains a specific frustration that gets blamed on the model. A long instruction, a long document and a final question produce a system that follows the instruction and answers about the end of the document, appearing to have skimmed. It did not skim; the middle of what it was given is the least dependable part, and the whole input was arranged so that the interesting material sat exactly there. The remedy is not a better model. It is deciding what actually needs to be in the input, and where in it those things go.BEGINNINGENDHandleddependablyReliabilityfallingLeastdependableReliabilityreturningHandleddependably
The shape is what makes this counterintuitive, because most degradation people meet is monotonic: things get worse the further you go, and you compensate by putting the important part first. Here the far end recovers, which means the intuition of front-loading everything is only half right and the material at genuine risk is whatever sits in the region nobody thinks about. It also explains a specific frustration that gets blamed on the model. A long instruction, a long document and a final question produce a system that follows the instruction and answers about the end of the document, appearing to have skimmed. It did not skim; the middle of what it was given is the least dependable part, and the whole input was arranged so that the interesting material sat exactly there. The remedy is not a better model. It is deciding what actually needs to be in the input, and where in it those things go.
03

Seen in the wild

  • A long document where a detail from the middle is left out of the summary.

    ChatGPT
  • Retrieving three relevant passages instead of sending an entire handbook.

    Glean
  • A long coding session where a constraint stated in the middle stops being applied.

    Claude Code
04

Common misconceptions

People assume

A bigger context window means better recall.

In fact

It means more can be sent. How dependably material is used is a separate property, and filling the larger window puts more of what you care about into the region where reliability is weakest.

People assume

If it missed something, it was not in the input.

In fact

It may have been present and passed over. Nothing reports that, so an omission looks exactly like an absence, and teams spend time checking whether material was included when it was.

05

Questions

How do we know whether this is affecting us?
Test it on your own material: put a distinctive fact in the middle of a long input, ask for it, and repeat with the same fact at the start. Comparing those two answers takes ten minutes and is more informative than any general claim.
Should we avoid long inputs altogether?
No, but treat filling the window as a decision rather than a default. Retrieving the passages that matter and sending only those is more reliable and cheaper at once, which makes it an easy trade wherever retrieval is possible at all.
What if everything genuinely has to be included?
Then use position deliberately. Put the most important material at the beginning or the end, and restate any critical instruction in both places. It costs nothing, it requires no tooling, and it works with the effect rather than hoping against it.
06

Key takeaways

  • Capacity is not recall: the published figure describes room, not care.
  • The middle is the least dependable position, and the effect grows with length.
  • An omission looks identical to an absence, so it is hard to diagnose.
  • Send less; where you cannot, put what matters at the ends.
08

Tools that use this

  • ChatGPT

    A detail from the middle left out of a summary.

  • Glean

    Three relevant passages instead of an entire handbook.

  • Claude Code

    A constraint stated mid-session that stops being applied.

Last checked August 2026

All glossary terms