Glossary category
Performance and cost
Speed, quality and what you pay for.
33 terms
A single point every AI request passes through, so keys, spending limits, logging and model choice are managed in one place.
Sending many requests together to be worked through when capacity allows, usually cheaper than asking for each one immediately.
A standard test used to score models against each other, useful for shortlisting and unreliable as a promise about your own work.
Allocating a shared tool's cost back to the teams using it, which changes behaviour faster than any usage policy.
An agreed minimum you undertake to spend over a period, usually traded for a better rate and worth sizing against realistic use.
How many requests a plan lets you run at the same time, usually the first limit a team meets when it moves past piloting.
The reason a long document costs more to ask about than a short one, because everything you send is charged for on the way in.
What one completed piece of work costs rather than what one person or one month costs, the unit that makes AI spending comparable to labour.
A prepaid unit a vendor charges against, worth checking carefully because one credit rarely means the same thing across two products.
A free tier is a usable product at no cost, with limits such as message caps or slower models; freemium describes the business model built around it, where the free version exists to sell the paid one.
Running a trained model to produce an answer, as opposed to training it in the first place.
What it costs to run a model each time it answers, which is why heavy use is metered and the strongest models are rationed.
The address your software sends requests to for a particular model, which is what you are really buying when you buy access.
The wait between sending a request and getting the answer.
Running a model on infrastructure someone provides, whether the model maker's, a cloud vendor's or your own.
Sending each request to whichever model suits it, keeping the expensive one for the hard cases and something cheaper for the rest.
What you are charged once an included allowance runs out, and the commonest unpleasant surprise on an AI invoice.
Paying for what you actually use rather than a flat subscription, which suits uneven demand and makes forecasting harder.
Reusing the work already done on a repeated piece of context, which cuts both cost and wait when the same material is sent often.
Paying to reserve a guaranteed amount of model capacity, which removes queueing at busy times and is charged whether you use it or not.
A cap on how many requests you may send in a given period, which protects the service and shapes how fast an automated job can run.
Paying a fixed amount per person per month, predictable to budget for and wasteful when only some of those people use it.
Model access with no capacity to manage, where you are charged per request and the provider handles the machines behind it.
A contractual promise about availability and response times, with agreed consequences, and the thing to ask for before depending on a tool.
A hard ceiling on what an account can consume, the control that turns an open-ended usage bill into a known maximum.
Returning an answer word by word as it is produced rather than all at once, which changes how fast a tool feels without changing how fast it is.
How much work a system gets through in a given time, the number that matters for bulk jobs where a single reply's speed does not.
The amount of use a plan includes before extra charges begin, the number worth finding before comparing two headline prices.
Charging by the amount of text going in and coming out rather than per person, which makes long documents cost more than short questions.
The share of time a service is actually available, usually published as a percentage and worth reading alongside what the vendor commits to contractually.
Paying for what you actually consume, which rewards light use and can surprise you in a busy month without a cap in place.
The ceiling a plan puts on how much you can do in a day, week or month before you are slowed, switched to a weaker model, or stopped.
A lower unit price at higher usage, which rewards consolidation onto one vendor and quietly raises the cost of leaving.