← All products

Ollama Model Selector

Ollama Model Selector cover

The cloud subscription hides the problem — for $20 a month.

Claude Pro runs $20 a month. Local AI runs free forever after the pick — the reusable cost is a single wrong choice: a model pulled for fame that doesn't fit your task, your VRAM, or how fast you need it. That mismatch, not the hardware, is why people try local AI once and walk away.

This selector turns picking into a decision you can make in 30 seconds: the three questions in order, honest 8 / 16 / 24GB tiers, and the quantization guide written for people, not for LLM metrics.

Stop downloading models that don't fit.

You don't need a bigger model. You need the RIGHT model for the task, your VRAM, and how fast you need it. This selector is the decision system — not a list.

What's inside

📄 The Selector (18-page PDF)

  • The Framework: task x hardware x speed — the 3 questions, in order
  • 3 VRAM Tiers: exactly what fits and runs on 8GB / 16GB / 24GB (honest — no overselling)
  • Task-Specific Matrix: coding / writing / reasoning / RAG / agents / vision -> the model for each
  • Decision Flowchart: 30 seconds from "what do I need?" to "pull this"
  • Quantization Guide: Q3 / Q4 / Q5 / Q6 — what to pick and when

⚙️ 10 Ready Modelfiles — writer, coder, analyst, chat, rag, agent, summariser, translator, vision, titles. Each: save it, ollama create -f Modelfile-, ollama run . All tuned for 8-16GB cards unless noted.

📊 XLSX Calculator — pick by task + VRAM; 10-modelfile index; quant reference.

🖨️ Cheatsheet (2-page PDF) — printable.

€12 once, not $20 a month

One payment, no subscription. Claude Pro is $20 a month for the hosted chat; the local model you pick runs on your own GPU at €0 a month. The honest reference: a free tool called llm-checker also matches models to hardware — use it. This selector earns its €12 by doing what that tool does not: it names the models for the tasks a freelancer actually does (writing, coding, RAG, agents, vision), speaks plain language, and ships the ready Modelfiles. It wins on VRAM-matching and on being non-geek — not on being the only option, because it isn't.

Your data stays on your machine

Ollama runs the model you pick locally on your own GPU. No cloud API key, no prompts leaving your drive, no $20-a-month chat subscription to rent.

One-time price · no subscription · editable · unwatermarked

€12 once. The PDF, the Modelfiles and the calculator are plain files you own. Nothing is paywalled, nothing is watermarked.

Who this is for

Anyone running Ollama who wastes time downloading the wrong models or fighting context-window truncation.

Who it's NOT for

People using only cloud APIs; those who want "one model to rule them all" (we recommend task-specific models).

Why trust this?

Every model here is one we actually run — tested on an RTX 3090 (24GB) and an 8GB laptop. Names verified July 2026 against Ollama's library. Unverified models are flagged, not guessed.