The situation

You found the model picker in your first week, set it to the strongest name on the list and never touched it again. You wanted the best answers. Since then a variable rename takes about as long as designing the data model did, and you have never opened the setting beside it, the one deciding how hard the model works whichever one you picked.

The idea in one paragraph

There are two dials here and most people find one. The first is which model: a vendor sells a family of them, differing by more than clever against simple, also by speed, price, how much text they hold at once and how recent their knowledge is. The second is effort: how much work the model spends on your request, set separately and applied to whichever model you chose. Depth and speed are both for sale, they are not the same purchase, and the expensive dial is not always the one that fixes your problem.

Two dials, and a third thing that is not one

What you turnWhat it changesWhen it is rightWhat it cannot fix
The modelCapability, speed, price, window size, knowledge ageHard work, or work bigger than a small windowA request that never said what finished means
The effortThinking time, tool calls, explanation before actingRoutine work you want quick, or one problem worth sitting withA model out of its depth
The requestNothing in the tool, everything in the taskA fluent, confident answer aimed elsewhereWork genuinely beyond the model

Only the first two cost money. The third is free and gets tried last.

How it actually works

A family is one vendor's several models sold at different points, described by tradeoff rather than rank: fastest with near-frontier intelligence, the best combination of speed and intelligence, for complex agentic coding. Down that list latency runs from fastest to slower and the top model's input price is ten times the bottom's. Two differences are not capability at all, and they catch people: the fastest member sees 200,000 tokens where the others see a million, and two models sold side by side have knowledge cutoffs more than a year apart (checked 10 August 2026).

Effort belongs to the request, not the model, and it is not a strict budget. One vendor calls it a behavioural signal: a lower setting thinks less or not at all on an easy input, makes fewer tool calls, and goes straight to the work instead of explaining first. That is a real saving, not just a shorter answer. Thinking is billed as output tokens even when it is never shown to you, so depth is bought at the most expensive rate there is. And the setting is written into the request, so changing it mid-conversation discards the cached prefix you were getting cheaply.

What this changes for you

  • Set the default low and reach up, not the reverse. One vendor's cost guidance names a usual cause of a surprising bill: the strongest model left on as the default for everything.
  • Try the effort dial before the model dial. It is the cheaper move and the one almost nobody finds, and on routine work it changes the waiting more than the answer.
  • Match the properties that are not capability. A job bigger than a small window needs the model with the larger one; a question about last spring's release needs the cutoff date, not a stronger model.

Where it breaks

More is not reliably better. Guidance on the highest effort setting says plainly that on structured or less demanding tasks it adds significant cost for small gains and can tip into overthinking. This dial has a wrong end at both ends.

Everything specific here expires. Names, tiers and prices move every few months, so a habit built on a model name outlives the model. Build it on the shape: cheap and fast at one end, deep and slow at the other, and a second dial travelling the same axis for less.

Some products decide for you. An automatic mode picks per request and removes the dial: you cannot match what you cannot set, and you are not told which one answered.

Terms used on this page

  • Model family — several models from one vendor at different points on capability, speed and price, sharing one interface, so switching is a setting not a migration.
  • Effort — how much work a model spends on one request: thinking time, tool calls, how much it explains. Set separately from the model.
  • Thinking tokens — the reasoning a model does before answering, billed as output whether or not any of it is shown to you.
  • Knowledge cutoff — the date past which a model has no reliable knowledge of the world.

Lesson 7 is not written yet. Until it is, the Level 1 path shows where this lesson sits, and the glossary collects the terms above.