Resources · Choosing the right AI
How much does it cost to self-host an LLM?
Updated 2026-09-15
The honest number first
A machine that runs a genuinely useful AI model, the kind that can read your documents and draft your emails all day, costs roughly $1,500 to $4,000 to put together. After that you pay for electricity, which for typical business use lands between $10 and $40 a month. The models themselves are free to download and licensed for business use.
That is the whole cash side of it. We run our own consulting work on exactly this kind of setup, a used workstation with one strong graphics card, so the numbers in this guide come from our own utility bills and our own receipts, not from a spec sheet.
The part that never makes it onto the invoice is the third cost, and it is the one that decides whether self-hosting makes sense for you. We will get to it.
What the money actually buys
One number controls everything: the memory on the graphics card. AI models have to fit inside that memory to run well, and a bigger model is a smarter model. So the card is not a nice-to-have, it is the whole purchase.
- Under $1,000, using hardware you have. A recent computer with a gaming graphics card runs small models. Good for experimenting and for narrow tasks like sorting and tagging text. Not close to ChatGPT for general work.
- $1,500 to $4,000, one dedicated machine. A workstation, often bought used, with a card carrying 24 to 32 GB of memory. This runs the mid-size open models that handle most back-office work well, for a whole small team, one request at a time.
- $8,000 and up, multiple cards. For running bigger models or serving many people at once. Very few small businesses need to start here, and you will know if you do because the single machine will tell you.
Prices on used workstation hardware move around, but the shape holds: the card is most of the budget, and memory on the card is what you are shopping for. If you want the specific shopping list, the cards, the box, and what to skip, see what on-prem AI hardware to actually buy.
The monthly bill
A machine like this draws a few hundred watts when it is working and much less when it sits idle. Run it through a normal business month at average U.S. commercial electricity rates and you get an electric bill in the range of ten to forty dollars. There is no per-seat license, no per-word charge, and no monthly subscription. Ask it a thousand questions, the bill does not change.
That flat cost is the quiet appeal of self-hosting. Cloud AI is a meter that runs every time anyone uses it. A machine in the closet is a fixed cost that gets cheaper per use the more your team leans on it.
Three ways to host a private LLM
“Private LLM” gets used for any model that runs where the public chat products cannot see your data. There are three places to host one, and they have very different bills.
- On your own hardware. The machine described above. Highest cost on day one, lowest running cost, and the only option where the data never leaves the building. Everything else in this guide assumes this path.
- On a rented GPU server. A cloud provider rents you a dedicated card by the hour, roughly one to three dollars depending on the card, and you run the same free models on it. Left on around the clock, that comes to roughly $700 to $2,200 a month, which is why renting is the right way to try before you buy and the wrong way to run full time. Your files go to that provider’s data center, so the model is private from the AI companies but not private from the host.
- As a private deployment inside a big cloud. Amazon, Microsoft, and Google will run open models, and in some cases the frontier models, inside your own cloud account, metered per use or by the provisioned hour. No hardware, no setup, and for light use it is the cheapest of the three. The trade is that you are back on a meter, and privacy rests on a contract and the provider’s settings rather than on a locked door.
If the reason you want a private model is a rule about where data may live, only the first path fully answers it. If the reason is cost control at volume, the first path wins on running cost and the second never does. If the reason is curiosity, start with the third or the second and buy nothing yet.
The cost nobody prices in
Somebody has to stand the machine up, choose and update the models, watch the temperatures, apply the updates, and notice when something quietly stops working. None of that is hard the way surgery is hard, but it is real recurring work, and it is the line item that sinks the naive math.
If you have a technical person who enjoys this, the time cost is small and the setup tends to keep getting better. If you do not, you will either pay someone to set it up and check on it, or the machine slowly becomes an expensive space heater. Be honest about which business you are before you buy anything.
When the API is cheaper, and when it is not
Here is the comparison that actually matters. A small business using cloud AI for ordinary work often spends somewhere between twenty and a couple hundred dollars a month on it. Against that, a $2,500 machine takes a year or two to pay for itself on subscription savings alone, and that is before counting anyone’s time.
So on pure dollars, light users should not self-host. The math flips in three situations:
- Your data cannot leave the building. Law firms, medical and dental offices, accountants. If the files legally or practically cannot go to a third party, the AI has to come to the data, and the hardware cost is just the price of admission. This is the strongest reason, and we cover it in frontier vs private on-prem AI.
- You use AI heavily, all day, every day. Constant document processing, always-on automation, work that would run a cloud meter around the clock. Fixed cost beats metered cost at volume.
- You need the cost to be predictable. A machine you own cannot surprise you in a busy month.
If none of those describe you, start with cloud models and revisit later. That is the cheaper path and there is no shame in it. Our guide on what AI costs a small business covers that side.
What this means for you
Budget $1,500 to $4,000 for the machine, a few hundred a year for power, and a real number of hours for care and feeding. Get the card with the most memory you can afford, because that decides what the machine can become as better free models keep arriving. And before any of it, make sure you actually have the reason: privacy, volume, or predictability. The businesses happiest with self-hosted AI bought it for one of those three, not to save twenty dollars a month.
If you want help sizing this honestly for your situation, including being told that you do not need it yet, that is exactly what the assessment is for.
Questions people ask
Can I run an LLM on a normal office computer?
Are the models free?
Is a private LLM the same as a self-hosted LLM?
What does it cost to host a private LLM in the cloud instead of buying hardware?
Is a self-hosted model as good as ChatGPT or Claude?
What about renting a cloud GPU instead of buying?
Sources
Want this figured out for your business?
The assessment tells you where AI is worth it for you, fixed price, and the fee comes off the build.
Get an assessment →