
This week I talked myself into switching models twice — once up, once down — and talked myself out of both by the time I'd actually run the numbers.
The setup I'm protecting is simple. My assistant runs on DeepSeek V4 Pro, the bigger of DeepSeek's two current models. It costs $0.435 per million input tokens and $0.87 per million output, a rate that's held since late May. It's a 1.6-trillion-parameter mixture-of-experts model, text-only, with a million-token context window. For months it's been the best cost-to-capability deal I've found for the reasoning and coding work my agent actually does.
The two temptations were opposite directions. Up was GPT-5.6 Terra. Down was DeepSeek V4 Flash. Both looked like obvious moves until I wrote out what I'd actually get.
The upgrade that isn't
GPT-5.6 Terra is OpenAI's "balanced" tier — the middle rung between the flagship Sol and the budget Luna. It went generally available in early July, and OpenAI cut the price at the end of July to $2 input and $12 output per million tokens. A million-token context, image input, thinking controls. On paper it's the grown-up version of everything I run today, with eyes.
The eyes are the real lure. V4 Pro is text-only. Terra can take an image. If I ever want my assistant to actually see a screenshot or a photo, Terra solves that in a way V4 Pro structurally can't.
But the cost is not subtle. On output — and reasoning models live and die on output, because thinking tokens bill as output — Terra runs about 14 times what V4 Pro charges. Fourteen dollars and change per million output tokens versus under a buck. My workloads are output-heavy, and high-thinking mode makes them more so.
And what does that 14x buy on the tasks I actually run? On the agentic and coding benchmarks, Terra edges out the April V4 Pro checkpoint. But DeepSeek shipped a refreshed V4 Pro this week, and by the independent leaderboards that refresh has mostly closed the gap. I'd be paying 14x for a model that, on my actual work, is roughly a draw with the newest thing I already have.
That's the whole case against the upgrade in one line: it's not that Terra is bad. It's that the premium buys a capability I don't use — vision — and a benchmark edge that a free update already erased.
The downgrade that isn't
The other direction was Flash. DeepSeek V4 Flash is the small sibling: 284 billion parameters, 13 billion active, MIT-licensed open weights, and $0.14 input / $0.28 output. About three times cheaper than Pro, and noticeably faster, with lower time-to-first-token.
For a while the pitch was almost irresistible: run Flash as the default, keep Pro for the hard stuff. The line I kept seeing in the community was that Flash at maximum thinking roughly matches Pro at high thinking. If that held, I'd cut my bill by two-thirds and barely notice.
It doesn't quite hold. The gaps between Flash and Pro are small on plain coding — a point or two on the standard benchmarks — but they're not small on the things I've learned to care about:
- Factual recall. On the harder factual benchmarks, Flash gets things wrong noticeably more often — call it roughly double the hallucination rate.
- Agentic work. On terminal-style and tool-use benchmarks, Flash trails Pro by a meaningful margin.
- Browsing and lookup. Pro is stronger at pulling apart web pages and staying on task.
For an assistant that's mostly chat, memory, and lookups, Flash is a genuinely good fit — cheap, fast, good enough. For the work my agent does — reasoning through code, driving tools, not inventing things — the cheaper model is the wrong place to save.
Staying put is still a decision
I expected this to feel like settling. It doesn't. There's a difference between refusing to change and deciding, with numbers in hand, that the current thing is still the right thing.
The decision I actually made is a routing one, not a loyalty one:
- V4 Pro stays the default for reasoning and coding, where output is already cheap and the capability ceiling matters.
- Terra gets used for exactly one job: when I need vision. That's a capability V4 Pro can't fake.
- Flash stays out of the main path, reconsidered only for light, high-volume, low-stakes work.
The price ladder, for the record:
- V4 Flash — $0.14 in / $0.28 out
- V4 Pro — $0.435 in / $0.87 out
- GPT-5.6 Terra — $2 in / $12 out
That's the shape of the modern model decision, and it's worth spelling out because it isn't the decision I used to make. I used to ask "which model is best?" Now I ask: which model is best for this one task at this price, and what does the next tier actually buy me?
One more wrinkle makes the exercise feel timely. DeepSeek has signaled that the flat rates I've been anchored to are about to change, moving toward peak and off-peak tiers. The exact numbers are still settling, but the era of "the cheap model is permanently cheap" looks like it's ending. Which means the routing decision I just made is probably the first of several — and getting comfortable re-running it is the actual skill here.
The expensive upgrade was a draw on my work. The cheap downgrade was a real cut in the places I care about. The middle model, for my workload, is still the answer.
How many of the model decisions you've been putting off would resolve themselves if you just wrote down the price and what the next tier actually buys?



