The Best Model Ever Is Too Dangerous to Use (Always)

The Benchmark Scroll: Proof That Proves Nothing
The Benchmark Scroll Is Not Your Workload

Every time a new frontier model is released, the model provider publishes a blog post about how it's the best model ever. And it is — the charts are clean, the bars are tall, the numbers are real. Every frontier model is better than every other frontier model. All the time. No model competes. All the models are the best model ever compared to all the other best models ever. It's a continuous spell.

What's crazy is that the benchmark scores keep getting closer to 100, but the frontier models keep one-upping each other. And when a benchmark maxes out — when every model is scoring 98 or 99 on the same test — they just find another benchmark. New test, new chart, new winner. The scroll never stops scrolling. And every model is cheaper than the last best model ever. Cheaper and better. The best model ever, and it costs less.

Now, the Benchmark Scroll doesn't tell you what the actual test was. You can go look up a benchmark, you can go see what the test was. But they're not sharing the details. What hardware did it use? What were the specific model settings? Where was it run? What was it compared to? Who ran it? Was the benchmark in the training data? Nobody is in a hurry to tell you. But more importantly — why are you so full of questions about the best model ever? Why can't you just accept that the new model is always going to be the best model ever? All the other models are the best model ever. The best model ever is the most important part of this.

The side effect of the Scroll is that anyone who reads any benchmark will comment that whatever model was just released is the best model ever. And it changes everything. The best model ever always changes everything. Because that's what the next model does — it changes everything. Because we're on the frontier. And on the frontier, every model is the best model ever.

Whether it's Anthropic, OpenAI, xAI, Z.ai, or DeepSeek — every time a new model is released, there is a company publishing a blog post about how it's the best model ever. And the cadence is every week. Every week there is another best model ever being released, even if it's not the best model ever. The announcement is always that it's better. It's better than the best model ever. It's cheaper than the best model ever. It changes everything, just like the last best model ever changed everything.

That's the Benchmark Scroll. It's a magical artifact, and it's constantly listening and adapting to questions. Any question you ask it, it's going to find a way to show you a bar chart that says the new model is the best model ever. It's sentient to "better."

The counter-move is not to ignore benchmarks. It's to run your own. Take the model and run it against your data, your architecture, your compliance constraints, and your team's actual skill level. If the model provider won't let you reproduce the benchmark on your environment, the benchmark is marketing.

But by the way — the best model ever? You'll only be able to use it for two weeks. Before the model provider announces it's too dangerous to use. Because it's also the most dangerous model. The most dangerous model is the best model ever. Just wait until what's next. What's next is another model that's going to be the best model ever, which is also going to be too dangerous to use.

The scroll never stops scrolling.


HQ 7 — Guided. Human led the creation, AI filled gaps and expanded points.


The Benchmark Scroll is one of the artifacts in The AI Developer's Field Guide, a field guide to the anti-patterns AI brings to software engineering.