AI models
The brains
The part that thinks. I have run more than ten of them, from small and quick to very large.
An AI model is a program that has read an enormous amount of text and can now read, write and reason. Its size is counted in “parameters”, written with a B for billions. Think of it as the number of adjustable connections in the brain. More usually means smarter, slower, and much hungrier for memory.
Most people who try this at home run models of 4 to 16 billion. My laptop has 128 GB of memory, which let me run models of up to 120 billion on my own desk, with nothing sent to anyone.
8 to 120
billion parameters: the range of models I have run on one laptop
The ones I have worked with
gpt-oss 120B
120 billion · localThe largest I run on my own machine. It takes about 63 GB of memory by itself. I use it for hard reasoning and for coding by hand.
Qwen3 Coder Next
80 billion · local · retiredA coding specialist. I retired it after real work showed it was the slowest of three candidates and its size bought nothing extra.
DeepSeek V4 Flash
about 91 GB on disk · local · retiredA very large open model. It ran, but stacked on everything else it exhausted the machine’s memory during testing. Too big and too slow for daily use, so it went.
Gemma 4, my own agent build
31 billion · localFor weeks the everyday brain of my assistant. It scored 100 out of 100 in a test shaped like real use, where the model before it scored 70. Today it is the one that can look at pictures.
Qwen3 and Qwen3 Coder
30 billion · local · retiredEarlier everyday brains and candidates. One answered with invented tool commands instead of really looking things up, which is how it lost its place.
“Libero”
27 billion · localA model with fewer built-in refusals, kept only on my own machine. It is switched on by request, for one conversation or a set time, and then the normal model comes back automatically.
gpt-oss 20B
20 billion · localThe everyday local brain since late August. Chosen after a head-to-head test on real tasks.
Qwen3 8B
8 billion · localThe quick one. It answers simple messages in about 0.7 seconds, where the bigger model needs 2.
An embedding model
localNot a talker. It turns text into numbers so that the archive can find notes by meaning, not just by matching words.
Claude, GPT and Gemini
from outside companiesMuch larger models that I reach over the internet for the hardest jobs and for a second opinion. My main chat assistant currently answers with one of these, with a local model as its backup.
What running them taught me
- 1
The biggest model is not the best one if it does not fit. A brain that pushes everything else out of memory helps nobody.
- 2
Choose by testing on real work. More than once the model that looked best on paper lost.
- 3
One brain cannot do every job. I tested using a single model for both quick and hard messages, and it was worse at both.