Glossary
On premises
Running software on hardware you control, in your own building or data centre, rather than reaching it as a service someone else operates.
In plain terms
The machines are yours and they are somewhere you control. Nothing is sent to a provider because the software runs where you put it. For AI that usually means buying capable hardware, running a model on it, and having somebody responsible for keeping all of that working, which is a different kind of commitment from a subscription.
Why it matters
Because it is the arrangement that answers the hardest version of every question in this category, and it answers them by making you responsible for everything instead. Where material genuinely cannot leave, that trade is correct and there is no substitute. Where it is chosen for general caution, the organisation has taken on a permanent operational burden to address a concern that a contract would have addressed.
How it works
It removes a category of question rather than answering it. Where nothing is sent, questions about retention, training, location, legal reach and who else touches the material do not arise, which is genuinely simplifying. That is the entire case for it, and it is a strong one where those questions are hard for you.
In exchange you take on everything a provider was doing. Hardware, capacity, updates, security patching, availability and somebody who understands all of it are now yours, and that person is the item most often left out of the comparison. The costs are real, continuing and mostly staffing rather than equipment.
The economics turn on how busy the hardware is, which is the calculation done most optimistically. Machines are paid for while idle, so a comparison against a per-use service has to assume realistic occupancy rather than peak. Steady heavy demand favours owning; spiky or occasional demand favours the meter by a wide margin.
The models available to run yourself are not the same set as the ones offered as a service. Open-weight models can be excellent and are a different shortlist, so the choice is not only about where something runs but about what you will be running. Teams comparing against a hosted product they already like should be explicit that they are comparing different things.
It is increasingly a deliberate exception rather than a default posture. Organisations that run most things as a service and one thing on their own hardware, because that one thing handles material that cannot leave, are making a narrower and more defensible choice than those treating it as a general stance. The narrow version also concentrates the operational cost where it buys something.
What moves when the hardware does
Seen in the wild
Running an open-weight model on hardware you own, where nothing is sent to any provider.
OllamaLoading and running models on a controlled machine through a desktop application rather than a command line.
LM StudioPutting a chat interface and document store in front of a locally run model so a team can use it like a product.
AnythingLLM
Common misconceptions
People assume
It is more secure by default.
In fact
It moves responsibility rather than raising a standard. Material stays under your control and patching, access management and availability become yours, so whether it is more secure depends on how well you would do those things compared with a provider whose specialists do only that.
People assume
It is cheaper at scale.
In fact
It can be, at high steady utilisation, and the comparison has to include idle time and the person maintaining it. Hardware is paid for whether or not anybody is using it, which is why spiky demand favours a meter and why the staffing line decides more of these comparisons than the equipment line.
Telling them apart
On premises vs Cloud
On premises
Your hardware, your responsibility, nothing sent anywhere.
Somebody else's hardware and responsibility, reached over the internet.
Ask who gets called when it stops working at seven in the morning.
Questions
- When is it the right answer?
- When material genuinely cannot leave the organisation, or when demand is steady and heavy enough that owning the hardware pays. Those two reasons cover most sound decisions. Choosing it as a general posture, rather than for one of them, takes on a permanent cost to address a concern a contract would have covered.
- What does it actually cost?
- Hardware, power and, mostly, somebody's ongoing attention over the whole life of the thing. The staffing line is the one left out of comparisons and the one that quietly decides them, because updates, patching, capacity and availability are all continuing work rather than a setup task that finishes.
- Do we get the same models?
- A different set. Open-weight models can be run yourself and can be very capable, and the newest hosted models are generally not among them. That makes this a choice about what you will be running as well as about where, which is worth making explicit when comparing.
- Is a hybrid arrangement sensible?
- Frequently it is the best answer available. Running most things as a service and one constrained workload on your own hardware concentrates the operational cost where it buys something, rather than applying the strictest arrangement to all of the work.
Key takeaways
- It removes a category of questions by making you responsible instead.
- The staffing cost decides these comparisons more often than the hardware cost.
- Idle hardware is still paid for, so spiky demand favours a service.
- You are choosing a different set of models, not only a different location.
- A narrow exception for constrained material beats a general posture.
Tools that use this
- Ollama
An open-weight model on hardware you own, sending nothing anywhere.
- LM Studio
The same arrangement through a desktop application.
- AnythingLLM
A usable interface and document store in front of a locally run model.
Last checked July 2026