> learn about ai / use it well
What a model is, and why there are so many
Big and small, fast and careful, open and closed. What the names on the menu mean and how to pick.
6 min read
The app is not the model
Open ChatGPT and you are looking at an app. The app has a chat box, a memory feature, a settings page. Behind it sits the model, the actual trained brain, and the app lets you pick which one. Claude is the same: the app is Claude, the models are Sonnet, Opus and Fable. Gemini is the same. Grok is the same.
This matters because the model decides how good the answers are, and the app decides what you can connect to. Two people can be using "ChatGPT" and getting very different results because one is on the free model and one is on the flagship.
Sizes
Every company sells a range, and the range follows the same shape:
Small and fast. Cheapest to run, answers in a blink, fine for simple tasks, weakest on hard reasoning. Anthropic's is Haiku. Google's is Flash. OpenAI's are the mini and nano versions.
Middle. The workhorse. Most people's daily model. Good at nearly everything, fast enough. Claude Sonnet, Gemini's main Flash tier these days, OpenAI's default model on paid plans.
Big and careful. The flagship. Slower, pricier, best at long, hard, multi-step work. Claude Opus and Fable, GPT-6, Gemini Pro.
For an agent that reads your inbox or watches a topic, the middle model is more than enough. For one that has to reason about money with hard rules, the bigger the better.
Versions
Names carry version numbers, and they matter. A model called 5 is a different generation from a model called 4, the way an iPhone is. Companies now ship a new generation roughly once a year and point releases every few months. If a guide online mentions a model name you cannot find in the menu, it has probably been replaced. That is normal.
Thinking mode
Most current models can "think" before they answer: work through a problem in steps, check themselves, then reply. It is slower and it costs more, and it is much better at maths, planning and anything with several moving parts. Apps usually give you a toggle, or pick it automatically. For a quick question, leave it off. For "plan my week around these constraints", turn it on.
Context window
Every model can hold a fixed amount in mind at once, measured in tokens (roughly three quarters of a word each). A few years ago that was a few pages. Today it is hundreds of pages to a million tokens, which is several novels. Two practical consequences:
- You can paste a whole contract or a year of statements and ask about it.
- A very long chat eventually fills the window and the model starts forgetting the beginning. When an agent "forgets its instructions", this is why. The fix is a fresh chat with the instructions at the top, or a Project that keeps them there.
Knowledge cutoff
A model stops learning on a date. Everything after that, it has never read. Ask it about last week's news without giving it a search tool and it will either say it does not know or, worse, guess. This is the single biggest reason to give an agent web search when its job involves anything current.
Open weights and closed
Some companies publish the trained model for anyone to download: Meta's Llama, DeepSeek, Mistral, Alibaba's Qwen. That is called open weights. Others keep the model on their own servers and sell access: Anthropic, OpenAI, Google. For a beginner the difference is mostly philosophical. Open models power a lot of products you use without knowing it, and they let a company run AI on its own machines. The best of them are close to the closed flagships and a few months behind.
So which one?
The honest answer for someone starting out: the default middle model of whichever app you already have is fine, and better than the flagship of two years ago. Pick the app by what you can connect to it and what your plan allows. Pick the model when you notice it struggling, and try the bigger one.
What to remember
- App and model are different things. The app decides connections, the model decides quality.
- Small, middle, big. Middle is the right default.
- Thinking mode for hard problems. Web search for anything after the cutoff.
- Long chats forget. Keep instructions where the app keeps them.
the one thing to do next
The fastest way to make this stick is to do one real thing with it today. Say who you are and we hand you one, prompt included.
Pick who you are