Compare

Ollama vs cloud LLM pricing.

Cloud vendors rent you AI by the token. Run the same models on your own hardware and the bill is the power draw. Here's the break-even math, with the honest caveats.

The 3-year math

€800 of hardware vs €200/month in subscriptions

Used RTX 3090
~€800

One-time. 24GB VRAM. Runs 7B–70B models at home.

Cloud LLM, 3 years
~€7,200

€200/mo × 36 months in token + subscription fees.

Kept in your pocket
~€6,400

Over three years, owning beats leasing, and your data never leaves the building.

Honest caveats

When cloud is genuinely better

  • Spiky, rare workloads: if you run AI twice a year, renting beats buying a GPU you'll underuse.
  • Frontier-only models: the largest cloud-only models still outclass local on some tasks. We say so.
  • Zero hardware tolerance: if a dead GPU means dead business, a managed provider's uptime has value.

For a solo operator running AI daily, local wins on cost and sovereignty. The Local AI Stack guide walks the exact setup; the Ollama Model Selector matches a model to your VRAM.