Skip to content

Glossary

Inference provider

The company that actually runs a model and serves its answers, which decides speed, availability and cost rather than capability.

In plain terms

Whoever owns the machines the model runs on when you press send. The same model served by two companies can feel quite different in speed and reliability, because those are properties of the running rather than of the model.

01

Why it matters

Because identical capability can arrive with quite different operational characteristics, and buyers frequently attribute those to the model. A tool that feels slow, or that fails at busy times, may be running a perfectly good model in an arrangement that cannot keep up, and that is a fixable problem of a completely different kind from a capability limit.

02

How it works

The same weights served by different operators produce different experiences. Speed, how many requests can run at once, what happens under load and how often something is unavailable are all decisions made by whoever is running it, not properties of the model itself.

Open models make this visible because anybody may serve them. A model available for others to run creates a market of operators competing on price, speed and reliability, so the choice of who runs it becomes a real decision rather than something bundled invisibly.

Where a model is closed, the two roles usually collapse into one company and the distinction stops being actionable. Knowing it is still useful, because it explains why some products can offer a choice of who serves their model and others structurally cannot.

Cost per request and speed pull against each other, and the operator sits at that dial. Serving faster generally means holding more capacity ready, which costs more whether or not anybody is using it, so a cheap arrangement and a fast one are usually different arrangements.

For a buyer the practical question is who to complain to when it is slow. Where the vendor runs their own serving they can fix it; where they buy it in, they are relaying a request to somebody else, and the difference shows up in how quickly a performance problem gets resolved rather than in any feature.

Two things a buyer experiences as one

Two things a buyer experiences as oneSeparating these is worth the effort because the misattribution is expensive. A team that finds a tool sluggish concludes the model is not good enough, starts evaluating alternatives, and spends weeks arriving at a different product that may well be served better and may not, without ever having identified what was actually wrong. The right-hand column has cheaper remedies: a different plan, a different operator for the same weights, or simply a vendor who can be told that performance is unacceptable. This also explains something buyers find puzzling, which is why the same named model appears to behave differently across products. The weights are identical and the experience is not, because holding capacity ready costs money whether or not anybody is using it, and every operator makes a different bet about how much to hold. That bet is invisible on a feature page and shows up on a Monday morning when everybody is working at once, which is precisely when a tool most needs to be dependable and least often is.Set by the modelWhat it can and cannot do.How good the answers are.What it knows about.Set by whoever runs itHow fast it replies.Whether it copes at busy times.What each request costs.A disappointing experience isusually attributed to theleft-hand column and isfrequently caused by the right.The two have different remedies,and only one of them requireschanging model.
Separating these is worth the effort because the misattribution is expensive. A team that finds a tool sluggish concludes the model is not good enough, starts evaluating alternatives, and spends weeks arriving at a different product that may well be served better and may not, without ever having identified what was actually wrong. The right-hand column has cheaper remedies: a different plan, a different operator for the same weights, or simply a vendor who can be told that performance is unacceptable. This also explains something buyers find puzzling, which is why the same named model appears to behave differently across products. The weights are identical and the experience is not, because holding capacity ready costs money whether or not anybody is using it, and every operator makes a different bet about how much to hold. That bet is invisible on a feature page and shows up on a Monday morning when everybody is working at once, which is precisely when a tool most needs to be dependable and least often is.
03

Seen in the wild

  • A tool that feels slow at busy times, which is a serving question rather than a limit of the model.

    ChatGPT
  • Choosing among several operators serving the same open model on price and speed.

    Hugging Face
  • Becoming your own operator by running the model on your own machine.

    LM Studio
04

Common misconceptions

People assume

A slow tool means a slow model.

In fact

Speed is mostly a property of how the model is served rather than of the model. The same weights run by a different operator, or with more capacity held ready, can feel entirely different while being the identical model.

People assume

This is the same as the company that made the model.

In fact

Sometimes, and not always. Where a model is open, anybody may run it, so the maker and the operator can be different companies with different reliability, different prices and different obligations to you.

05

Questions

Why does the same model feel different in two products?
Because serving decisions differ. How much capacity is held ready, how requests are queued and what happens at peak times are all chosen by whoever runs it, and those choices produce most of the difference a user notices between two tools using identical weights.
Does this matter if we just buy a finished product?
It matters when performance disappoints, because it decides who can fix it. A vendor running their own serving can act; one buying it in is passing your complaint along, and that difference shows up in resolution time rather than in anything visible during a trial.
Is running it ourselves cheaper?
It moves the cost from a per-request charge to hardware and people, which changes the shape rather than reliably reducing the total. It becomes attractive at steady high volume and is usually poor value for occasional use, because idle capacity is still paid for.
06

Key takeaways

  • Speed and availability are serving decisions, not model properties.
  • Open models create a market of operators; closed ones usually do not.
  • Cheap and fast are generally different arrangements, not one choice.
  • Who runs it decides who can fix it when performance disappoints.
08

Tools that use this

  • ChatGPT

    Slowness at busy times is a serving question, not a model limit.

  • Hugging Face

    Choosing among operators serving the same open model.

  • LM Studio

    Running it yourself makes you the operator.

Last checked August 2026

All glossary terms